gsoc-2026
GSoC 2026 Final Week
Coding wrapped in Week 14, but I had two things still open: the extraction-framework bugs
I filed as issue #845 needed an actual fix, and agentic-amdbpedia needed to go
from a working pipeline to something a mentor could sit down and use end to end. This week
I closed both out.
F Aug 28, 2026 – Sep 4, 2026
Two things carried over from Week 14: the extraction-framework bugs I filed as issue #845
still needed an actual fix, and agentic-amdbpedia still needed to go from a
working pipeline to something a mentor could sit down and use end to end. This week I
finished both.
Fixing what issue #845 found
Issue #845, which I filed against dbpedia/extraction-framework on August 27,
documents seven distinct bugs I found auditing the amwiki-20260801 dump: four
in EthiopianDateParser and its config, three in shared infrastructure that
affects every language, not just Amharic. Each one shipped with dump-level evidence --
exact article examples, occurrence counts, live triples on am.dbpedia.org --
rather than just a description of the symptom.
PR #846 fixes all seven of them: 231 additions and 16 deletions across 11 files, four commits, closing #845 on merge.
| Bug | Measured impact | Fix |
|---|---|---|
| A — Ethiopian→Gregorian year off by one | 9 spot-checked dates, all wrong by exactly ±1 year | Branch on the Gregorian month the JDN algorithm actually needs, not the Ethiopian one |
| B — Standard ቀን date format never parsed | 94.5% of Ethiopian dates in the dump silently dropped | Added a regex admitting ቀን ("day") between the day and year |
| C — Two date regexes could never match | dateRegex2/3 declared 4 capture groups, destructured into 3 | Fixed the arity mismatch so all five regexes actually run |
| D — ~9% of month spellings unrecognised | Month 4: 1,193 of 1,195 uses (99.8%) used an unlisted spelling | Extended monthsMap with the spellings actually seen in the dump |
| E — ignoreProperties overrides, doesn't union | 2,888 image/map/flag triples (7% of the dataset) leaked through | Language-specific ignore list now unions with the English baseline |
| F — Empty literals emitted as facts | 902 blank-string triples, 701 of them in mappingbased-literals | Empty parses are dropped instead of written out as ""@am |
| G — Stale Module namespace mapping | 39 ሞጁል: pages misclassified as articles, plus NPEs | Recognises the localised ሞጁል namespace while keeping legacy aliases |
Bug B affected the most dates: 94.5% of Ethiopian dates in the dump used the standard <month> <day> ቀን <year> form, and none of the five date
regexes admitted it, so almost every date was silently dropped rather than mis-parsed. Bug
E affected the most triples: because the Amharic ignore-list replaced the English baseline
instead of adding to it, 7% of the entire infobox-properties dataset was image filenames and
flag references emitted as facts.
I validated the PR the same way I wrote the bug report -- against real data, not just unit
tests in isolation: Java 8 core test compilation (BUILD SUCCESS), a focused
JUnit regression suite (15 tests, all passing), and a full Java 8 reactor build through the server module. It's open and awaiting maintainer review as of this writing.
agentic-amdbpedia: the pipeline, end to end
I renamed the project from cross-lingual-knowledge-assistant to agentic-amdbpedia this stretch -- a better name for what it actually is now: a
LangGraph-orchestrated pipeline that takes an Amharic Wikipedia infobox, retrieves
candidate DBpedia ontology properties with a hybrid Afro-XLM-R dense + BM25 sparse
retriever, reranks them with a Gemma 2 9B predictor, and puts the result in front of a
human before anything touches the live wiki.
The retriever and predictor aren't new models picked in isolation -- they're the same ones benchmarked in LLMIntegration's Issue #3: Afro-XLM-R as the dense retrieval backbone, and Gemma 2 9B -- evaluated there as a multilingual capability probe -- carried over as the production reranking/prediction model. Reusing the benchmarked pair instead of picking new models means the pipeline's accuracy numbers trace back to real Week 11 measurements rather than an untested guess. Candidate properties are also checked against the same shared mappings spreadsheet used throughout the summer, so the agentic pipeline's output stays consistent with the manually-tracked mapping work instead of duplicating it in a separate system.
Every accepted or rejected mapping is captured into an append-only training-example log, written in the same format LLMIntegration uses for its own benchmark datasets -- so every real decision a reviewer makes becomes more labeled data for the next round of experiments, not a one-off click. Publishing back to the wiki is consent-gated through MediaWiki bot credentials, and before I trusted any of it, I loaded the extracted output into Tentris and checked with real SPARQL queries, not just spot-read by eye.
Getting there also meant paying down infrastructure debt that a working local .env had been masking: CI was quietly broken because constructing the app at
import time forced a GROQ_API_KEY the CI environment never had, the
integration test that expects a real Postgres instance never actually had one in CI, and an
unpinned mcp dependency had been silently failing the scheduled end-to-end job
for weeks on an unrelated breaking release. I fixed all three and they're green now,
alongside a disk-cached retrieval index so the pipeline no longer rebuilds its embeddings
on every cold start.
Where things stand
PR #846 is the upstream outcome of the summer's extraction-framework audit: seven bugs,
evidenced with real dump data, now fixed and validated, waiting on maintainer review. agentic-amdbpedia is the working-tool outcome: a pipeline that goes from a
pasted infobox or a bare Wikipedia link to a reviewed, published DBpedia mapping, with every
decision I log for the next round of benchmarking back in LLMIntegration.