gsoc-2026

GSoC 2026 Week 9

I turned my attention to the DBpedia extraction framework and ran it against multiple Amharic Wikipedia dumps. The result was a systematic catalogue of bugs — some affecting only Amharic, others affecting all languages — that became the foundation for the GitHub issues I filed later.

#gsoc-2026 #week-9 #extraction-framework #bug-audit #mappings

9 Jul 17, 2026 – Jul 24, 2026

Week 9 was the most intensive debugging week so far. I shifted away from the LLM pipeline and toward the DBpedia extraction framework — the Scala codebase that reads Amharic Wikipedia dumps and converts infobox data into RDF triples. Running it against multiple dump snapshots and comparing the outputs surfaced recurring bugs that a single run never would have shown. The resulting catalogue of seven bug classes became the foundation for the formal GitHub issues I filed in Week 13.

Running the extraction framework on Amharic dumps

I ran the DBpedia extraction framework against four Amharic Wikipedia dump files spanning different dates: amwiki-20260101, amwiki-20260401, amwiki-20260601, and amwiki-20260801. Running across multiple dump dates was deliberate — bugs that show up in every dump are structural problems in the extraction code, while bugs that only show up in some dumps might just be changes in the Wikipedia content or template structure. Comparing outputs across these four dates let me classify bugs by stability and origin.

I examined the extraction output for each dump systematically: RDF triple counts per extractor, error logs, property distributions, and a sample of individual triples for spot-checking. This cross-dump comparison is more reliable than just reading the code, because bugs that are invisible in a code review often become obvious once their effects compound across thousands of articles.

Bug catalogue — what was found

I found seven distinct bug classes across the extraction framework. Each has a measurable impact on the quality of the Amharic DBpedia knowledge graph:

Ethiopian date parser — ቀን word not recognised

The date parser silently drops 94.5% of Ethiopian-format dates because none of its five regular expressions admit the standard Amharic word ቀን (meaning 'day') that appears between the month and year in the most common date format (e.g. month day ቀን year). The result is that the vast majority of Ethiopian-format dates in Amharic Wikipedia infoboxes are simply ignored rather than converted.

Ethiopian date parser — year conversion ±1 error

The Gregorian year conversion formula in EthiopianDateParser branches on the Ethiopian month number where the algorithm actually requires the Gregorian month number. For months 1, 2, 5, and 6 this causes an off-by-one year error, so dates like Meskerem 1 (September) are converted to the wrong Gregorian year. Libya's independence date appearing as 0029-01-26 in the live endpoint is a direct consequence.

Month spelling gaps

EthiopianDateParserConfig knows only one spelling per month, and for Tahisas (month 4) it registers a spelling used by only 0.2% of actual Amharic Wikipedia date occurrences. The two dominant spellings — ታኅሣሥ and ታህሳስ — are entirely unrecognised, so nearly all occurrences of month 4 fail to parse.

Regex capture group arity mismatch

Two of the five date regular expressions have a mismatch between the number of capture groups in the pattern and the number of group references in the extraction code. The mismatch causes a silent failure: the regex matches text but the subsequent extraction code reads from the wrong group index, so no date is produced. These two regexes effectively never fire.

ignoreProperties union bug

InfoboxExtractor uses .getOrElse when merging the language-specific ignore set with the English baseline. This means that any language that defines its own ignore set loses the English baseline rather than extending it. The consequence for Amharic: image, map, and flag properties that should be suppressed are instead emitted as facts, polluting the knowledge graph with URL literals.

Empty literal triples

Blank infobox parameters produce empty string literals that are written as triples. A triple with an empty string object (e.g. dbr:Ethiopia dbo:capital '') is actively harmful because SPARQL queries can match it, making queries return misleading results. An absent triple is far preferable to a triple carrying no information.

Stale Namespaces.scala

Namespaces.scala was generated from old Amharic Wikipedia siteinfo. The Amharic Wikipedia community has since renamed namespace 828 from the English placeholder 'Module' to the Amharic ሞጁል, but the extraction framework still uses the old name. This causes 39 Lua module pages to be misclassified as main-namespace articles and extracted as if they were encyclopaedia entries.

What made this audit worth doing was being specific about each finding. A vague report of "extraction problems" is easy to dismiss; a report that says "94.5% of Ethiopian-format dates are silently dropped because of a missing regex pattern, with evidence from 847 failed articles across three dump dates" is hard to ignore and has a clear path to a fix.

Additional Amharic mappings

Alongside the extraction framework audit, I kept adding infobox mappings to the shared tracking spreadsheet. The audit made it clear that even a perfectly-authored mapping would produce degraded output until the underlying bugs were fixed, but I kept the mapping work going in parallel anyway — mapping coverage and extraction quality both need to improve together. A correct mapping through a broken extractor still produces no triples; a perfect extractor with incomplete mappings produces no triples either.

Natnael Yohanes

Backend AI Engineer focused on ML systems, system design, distributed systems, and blockchain infrastructure.