gsoc-2026
GSoC 2026 Week 12
I got the baseline benchmark PR merged with full documented results, started Issue #4 — a systematic prompt engineering experiment — and confirmed the Ethiopian calendar extraction bugs from Week 9 with real extraction output evidence.
12 Aug 7, 2026 – Aug 14, 2026
Week 12 was about consolidating and moving forward. Mentors reviewed the Issue #3 PR, I merged it into main, and documented its results in the project's progress report. The next experiment — Issue #4, on how different prompt engineering strategies affect accuracy — started right away. And a parallel thread of work confirmed, with real extraction evidence rather than just code inspection, that the Ethiopian calendar bugs from Week 9 were causing real data loss in the Amharic DBpedia graph.
Issue #3 PR: merging the baseline
The pull request for Issue #3 was the first formal code contribution to LLMIntegration that
went through a full review cycle. It included the benchmark harness (configurable via
YAML), the Afro-XLM-R retrieval index builder, the DSPy-based reranker, the answer snapping
logic, and the results CSV for every evaluated model and shot setting. Mentor feedback
focused on two things: the reproducibility of seed sampling (fixed via random.Random(SEED + seed)) and how clearly the accuracy metric was defined
(I documented the full fuzzy-match formula in the README).
The merged results showed the best baseline configuration — Afro-XLM-R retriever feeding 5-shot Qwen 2.5 (7B) — hit about 62.4% accuracy on the 279-example test set. That's a real improvement over the retrieval-only floor of ~52%, but it also made clear there was still a lot of room for improvement through better prompt design, cross-lingual reasoning, and ensemble methods — which is exactly what Issues #4, #5, and #6 set out to do.
Issue #4: prompt engineering experiment
The question behind Issue #4: does how I instruct the LLM to reason about the shortlist of candidates affect accuracy, and if so, which strategy works best for Amharic property mapping? I kept the retrieval stage fixed (Afro-XLM-R, top 10) and only varied the prompt structure and reasoning pattern.
Zero-shot (baseline)
No examples provided. The LLM must rely entirely on its pre-trained knowledge of English DBpedia properties and its ability to interpret Amharic text.
Result: Accuracy ~58–62%; strong for common properties, weak for rare ontology classes.
Few-shot with LabeledFewShot
DSPy's LabeledFewShot optimizer selects k demonstrations from the training set that most closely resemble the test example, sampling with a fixed random seed for reproducibility.
Result: Best at k=5: +4.2 pp over zero-shot on average across models.
ChainOfThought (CoT)
The DSPy ChainOfThought module prepends 'Let's think step by step' reasoning to the prediction. The model describes why each candidate does or doesn't match before committing to a choice.
Result: Modest improvement over standard few-shot; reasoning traces helpful for debugging failures.
Self-consistency ensemble
Three independent CoT samples are taken for each input, and the majority vote is used as the final answer. Tied votes are broken by retriever rank order.
Result: Small but consistent gain over single-sample CoT; adds latency per example.
I used DSPy's module system throughout, implementing each strategy as a composable DSPy module instead of a hand-crafted prompt string, which made it easy to swap strategies in and out without rewriting the evaluation loop. The results confirmed few-shot examples and chain-of-thought reasoning both help, but neither alone closes the gap to the theoretical ceiling — which is what motivated the translation and ensemble experiments in Weeks 13–14.
Ethiopian calendar bugs: extraction evidence
I'd found the Week 9 bugs mostly through code inspection: reading the extraction framework source and spotting the missing regex, the wrong offset, the absent month names. This week I got something more concrete — actual extraction output showing these bugs causing real data loss at scale.
Year-off-by-one for months 1, 2, 5, 6
The Ethiopian calendar conversion code uses a fixed Julian Day Number offset, but months at the turn of the Ethiopian year (Meskerem and Tikimt) and the middle of the year (Ginbot and Sene) are off by ±1 Gregorian year depending on the leap year cycle.
Missing ቀን regex (94.5% of dates dropped)
The word ቀን (meaning 'day') appears in most Ethiopian date strings but was absent from the extraction regex. Without it, 94.5% of date literals in the Amharic dump were silently discarded rather than extracted.
Month 4 (Hamle) nearly unrecognised
Month 4 — ሐምሌ — had a spelling variant that appeared in 99.8% of Amharic Wikipedia articles but was not in the month-name lookup table. The result: virtually no July dates extracted from Amharic Wikipedia.
The most striking finding was the missing ቀን regex: because a word meaning "day" wasn't in the pattern, the extraction framework was silently discarding 94.5% of date literals from the Amharic dump. That's not a minor edge case — it means almost every Ethiopian-calendar date in every Amharic Wikipedia article had been silently dropped instead of extracted. This confirmed the extraction framework needed real upstream fixes, not workarounds in the mapping layer, so I started preparing formal bug reports for the following week.