gsoc-2026

GSoC 2026 Final Week

Coding wrapped in Week 14, but I had two things still open: the extraction-framework bugs I filed as issue #845 needed an actual fix, and agentic-amdbpedia needed to go from a working pipeline to something a mentor could sit down and use end to end. This week I closed both out.

#gsoc-2026 #final-week #extraction-framework #agentic-amdbpedia

F Aug 28, 2026 – Sep 4, 2026

Two things carried over from Week 14: the extraction-framework bugs I filed as issue #845 still needed an actual fix, and agentic-amdbpedia still needed to go from a working pipeline to something a mentor could sit down and use end to end. This week I finished both.

Fixing what issue #845 found

Issue #845, which I filed against dbpedia/extraction-framework on August 27, documents seven distinct bugs I found auditing the amwiki-20260801 dump: four in EthiopianDateParser and its config, three in shared infrastructure that affects every language, not just Amharic. Each one shipped with dump-level evidence -- exact article examples, occurrence counts, live triples on am.dbpedia.org -- rather than just a description of the symptom.

PR #846 fixes all seven of them: 231 additions and 16 deletions across 11 files, four commits, closing #845 on merge.

BugMeasured impactFix
A — Ethiopian→Gregorian year off by one9 spot-checked dates, all wrong by exactly ±1 yearBranch on the Gregorian month the JDN algorithm actually needs, not the Ethiopian one
B — Standard ቀን date format never parsed94.5% of Ethiopian dates in the dump silently droppedAdded a regex admitting ቀን ("day") between the day and year
C — Two date regexes could never matchdateRegex2/3 declared 4 capture groups, destructured into 3Fixed the arity mismatch so all five regexes actually run
D — ~9% of month spellings unrecognisedMonth 4: 1,193 of 1,195 uses (99.8%) used an unlisted spellingExtended monthsMap with the spellings actually seen in the dump
E — ignoreProperties overrides, doesn't union2,888 image/map/flag triples (7% of the dataset) leaked throughLanguage-specific ignore list now unions with the English baseline
F — Empty literals emitted as facts902 blank-string triples, 701 of them in mappingbased-literalsEmpty parses are dropped instead of written out as ""@am
G — Stale Module namespace mapping39 ሞጁል: pages misclassified as articles, plus NPEsRecognises the localised ሞጁል namespace while keeping legacy aliases

Bug B affected the most dates: 94.5% of Ethiopian dates in the dump used the standard <month> <day> ቀን <year> form, and none of the five date regexes admitted it, so almost every date was silently dropped rather than mis-parsed. Bug E affected the most triples: because the Amharic ignore-list replaced the English baseline instead of adding to it, 7% of the entire infobox-properties dataset was image filenames and flag references emitted as facts.

I validated the PR the same way I wrote the bug report -- against real data, not just unit tests in isolation: Java 8 core test compilation (BUILD SUCCESS), a focused JUnit regression suite (15 tests, all passing), and a full Java 8 reactor build through the server module. It's open and awaiting maintainer review as of this writing.

agentic-amdbpedia: the pipeline, end to end

I renamed the project from cross-lingual-knowledge-assistant to agentic-amdbpedia this stretch -- a better name for what it actually is now: a LangGraph-orchestrated pipeline that takes an Amharic Wikipedia infobox, retrieves candidate DBpedia ontology properties with a hybrid Afro-XLM-R dense + BM25 sparse retriever, reranks them with a Gemma 2 9B predictor, and puts the result in front of a human before anything touches the live wiki.

The retriever and predictor aren't new models picked in isolation -- they're the same ones benchmarked in LLMIntegration's Issue #3: Afro-XLM-R as the dense retrieval backbone, and Gemma 2 9B -- evaluated there as a multilingual capability probe -- carried over as the production reranking/prediction model. Reusing the benchmarked pair instead of picking new models means the pipeline's accuracy numbers trace back to real Week 11 measurements rather than an untested guess. Candidate properties are also checked against the same shared mappings spreadsheet used throughout the summer, so the agentic pipeline's output stays consistent with the manually-tracked mapping work instead of duplicating it in a separate system.

Every accepted or rejected mapping is captured into an append-only training-example log, written in the same format LLMIntegration uses for its own benchmark datasets -- so every real decision a reviewer makes becomes more labeled data for the next round of experiments, not a one-off click. Publishing back to the wiki is consent-gated through MediaWiki bot credentials, and before I trusted any of it, I loaded the extracted output into Tentris and checked with real SPARQL queries, not just spot-read by eye.

Getting there also meant paying down infrastructure debt that a working local .env had been masking: CI was quietly broken because constructing the app at import time forced a GROQ_API_KEY the CI environment never had, the integration test that expects a real Postgres instance never actually had one in CI, and an unpinned mcp dependency had been silently failing the scheduled end-to-end job for weeks on an unrelated breaking release. I fixed all three and they're green now, alongside a disk-cached retrieval index so the pipeline no longer rebuilds its embeddings on every cold start.

Where things stand

PR #846 is the upstream outcome of the summer's extraction-framework audit: seven bugs, evidenced with real dump data, now fixed and validated, waiting on maintainer review. agentic-amdbpedia is the working-tool outcome: a pipeline that goes from a pasted infobox or a bare Wikipedia link to a reviewed, published DBpedia mapping, with every decision I log for the next round of benchmarking back in LLMIntegration.

Natnael Yohanes

Backend AI Engineer focused on ML systems, system design, distributed systems, and blockchain infrastructure.