gsoc-2026
GSoC 2026 Week 13
I submitted Issue #4 (prompt engineering) as PR #9 and Issue #5 (translation experiment) as PR #10, formally reported all seven extraction framework bugs, and wrote an integration approach document laying out how the LLM pipeline and the extraction framework would eventually work together.
13 Aug 14, 2026 – Aug 21, 2026
Week 13 was one of the most productive of the summer. I submitted two pull requests in a single week — PR #9 closing Issue #4 (prompt engineering) and PR #10 closing Issue #5 (cross-lingual translation experiment). Alongside that, I wrote up the seven extraction framework bugs documented over Weeks 9 and 12 as formal GitHub issues on the upstream repository, and drafted an integration approach document describing how the LLM mapping pipeline would connect to the broader DBpedia extraction workflow once the individual experiments were done.
PR #9: Issue #4 prompt engineering results
The prompt engineering experiment gave me clear guidance on which strategies are actually worth using for Amharic property mapping. The pull request included the full experiment code, all result CSVs, and a markdown summary of findings. Chain-of-thought reasoning produced about a 2.1 percentage point improvement over standard few-shot across all models, with the biggest gains on examples where the correct property isn't the most surface-similar candidate in the shortlist — cases where reasoning helps the model look past misleading synonyms.
The self-consistency ensemble (three CoT samples, majority vote) added another 1.3 percentage points on top of single-sample CoT, at roughly 3× the latency per example — acceptable for an offline mapping pipeline, but it would need addressing for any real-time use case. These findings fed directly into Issue #6 (ensemble methods), which explored more principled ensemble strategies.
PR #10: Issue #5 translation experiment
The hypothesis behind Issue #5 was that translating the Amharic property mention to English before handing it to the LLM might improve accuracy, since the LLM has much richer English-language training signal for DBpedia property names than it does for Amharic. I evaluated two translation strategies:
Pivot translation (Amharic → English)
The Amharic property mention is translated to English before being passed to the retriever and LLM. This allows the LLM to operate entirely in English, where it has much stronger priors about DBpedia property names.
Pros
Leverages LLM's English-language knowledge; works with any English-only retriever.
Cons
Translation errors compound with prediction errors; loses Ge'ez script context the LLM might use.
Augmented retrieval (Amharic + English)
The retriever embeds both the original Amharic mention and its English translation, then takes the union of the top-5 candidates from each query. The LLM reranks the combined shortlist of up to 10 unique candidates.
Pros
More robust to translation errors; keeps both language signals alive for the LLM.
Cons
Larger shortlist can dilute the most relevant candidate; slightly higher latency.
I used translation models from the Helsinki-NLP Opus-MT family, which provide open-source Amharic-to-English translation. The augmented retrieval approach (using both Amharic and English queries) beat pivot translation alone, giving a net improvement of about 2.8 percentage points over the best baseline from Issue #3 — confirming that Amharic context carries signal that shouldn't be discarded even when English translation is available.
Extraction Framework: all seven bugs formally reported
Over Weeks 9 and 12 I'd accumulated evidence for seven distinct bug classes in the DBpedia extraction framework's Amharic language support. This week I wrote those findings up as formal GitHub issues on the upstream extraction-framework repository, with reproduction steps, expected vs. actual output, and impact estimates.
Ethiopian calendar date parser (±1 year off)
Months 1, 2, 5, and 6 of the Ethiopian calendar year produce incorrect Gregorian year conversions due to a fixed Julian Day Number offset that does not account for leap year boundary crossings.
Missing ቀን regex — 94.5% date loss
The word ቀን ('day') is present in virtually every Ethiopian date string but was absent from the extraction regex, causing 94.5% of date literals to be silently discarded.
Capture-group arity mismatch
A regex used for parsing date components declares more capture groups than the post-processing code expects, leading to an ArrayIndexOutOfBoundsException on certain inputs.
Month 4 (Hamle) 99.8% unrecognised
The spelling variant of ሐምሌ (Hamle, the fourth month) used in almost all Amharic Wikipedia articles was missing from the month-name lookup table.
ignoreProperties union bug
When two mapping files both list the same property in their ignoreProperties set, the union operation silently overwrites one entry with the other instead of merging them.
Empty literal emission
Certain template fields with whitespace-only values are extracted as empty string literals rather than being skipped, producing invalid RDF triples.
Stale Namespaces.scala
The namespace constants file references a DBpedia namespace URI that was retired, causing some triples to be serialised with an outdated prefix.
Integration approach document
The LLM benchmarking work and the extraction framework work had been running as parallel tracks all summer. This week I wrote a document connecting the two: an integration approach describing how the property mapping pipeline would work within the larger DBpedia extraction workflow once it's production-ready.
The integration I proposed works like this: the extraction framework runs its existing heuristic extraction pass over the Amharic Wikipedia dump, producing candidate property assignments for each entity. For assignments where the heuristic confidence is below a threshold, the LLM mapping pipeline runs as a second pass, using the Amharic infobox field as the query and the Afro-XLM-R retriever to surface candidates. The LLM picks from the shortlist and its prediction replaces the low-confidence heuristic assignment. High-confidence heuristic assignments pass through unchanged, so the LLM pass only costs compute proportional to the fraction of ambiguous cases.
I shared this document with mentors, and it became the basis for the final week's progress report and presentation.