gsoc-2026

GSoC 2026 Week 10

With the extraction framework bugs catalogued in Week 9, I shifted to a different kind of ground truth this week: querying the live Amharic DBpedia graph directly via SPARQL to see what data had actually made it through extraction, and refactoring Amharic Wikipedia articles to improve what would make it through in future runs.

#gsoc-2026 #week-10 #sparql #wikipedia #mappings

10 Jul 24, 2026 – Jul 31, 2026

After Week 9's deep audit turned up seven bug classes in the DBpedia extraction framework, I took a step back from the pipeline itself this week and looked at what it had already produced. Using SPARQL queries against the Amharic DBpedia endpoint, I checked the coverage and quality of the existing structured data — which entity types were well-represented, which properties were actually populated, where the gaps were biggest. Alongside that query work, I spent time refactoring eight Amharic Wikipedia articles to make sure they used the correct infobox templates and property names, since that directly determines what the extraction framework can pull out of the next dump.

Why SPARQL?

The extraction framework produces RDF triples, and those triples get loaded into a SPARQL endpoint where any query client can ask structured questions about the data. SPARQL — the W3C standard query language for RDF graphs — lets you ask things like "how many Amharic-labelled entities have a dbo:birthDate value?" or "which DBpedia ontology properties appear most frequently across all Amharic entities?" and get exact counts from the live graph instead of estimates from a sample.

This mattered because the LLM benchmark experiments in Weeks 11–13 would measure accuracy on a fixed test set of 279 examples. Knowing the underlying coverage of the Amharic DBpedia graph — which entity types were dense with data and which were sparse — helped me calibrate what accuracy levels were actually achievable given the data that exists. A pipeline predicting properties that rarely appear in the graph has no way to be verified empirically even when the LLM prediction is correct.

There was also a practical access problem: the main Amharic DBpedia SPARQL endpoint sits behind a VPN that needs institutional credentials to reach. This is what eventually led to the kgproxy solution I deployed in Week 14 — a lightweight AWS proxy that forwards authenticated SPARQL traffic from the public internet to the internal endpoint. For now I just ran queries from behind the VPN, but the friction was real and worth documenting.

SPARQL queries used this week

Here are the four key queries I wrote and ran this week, each serving a different diagnostic purpose, with the reasoning behind each — the queries are only useful in context.

Count Amharic entities by DBpedia class

This was the first diagnostic query — establishing how many entities of each type the Amharic DBpedia currently knows about. The results showed that Person and Place entities dominated, while Organisation and Work entities were sparse.

PREFIX dbo: <http://dbpedia.org/ontology/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
PREFIX rdf:  <http://www.w3.org/1999/02/22-rdf-syntax-ns#>

SELECT ?class (COUNT(DISTINCT ?entity) AS ?count)
WHERE {
  ?entity rdf:type ?class ;
          rdfs:label ?label .
  FILTER(LANG(?label) = "am")
  FILTER(STRSTARTS(STR(?class), STR(dbo:)))
}
GROUP BY ?class
ORDER BY DESC(?count)
LIMIT 20

Fetch Amharic entity properties and values

Used to pull a sample of property–value pairs for Amharic-labelled entities so we could verify which DBpedia ontology properties were actually populated versus which ones existed in the mapping definitions but had no data.

PREFIX dbo: <http://dbpedia.org/ontology/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>

SELECT ?entity ?label ?property ?value
WHERE {
  ?entity rdfs:label ?label .
  ?entity ?property ?value .
  FILTER(LANG(?label) = "am")
  FILTER(STRSTARTS(STR(?property), STR(dbo:)))
}
ORDER BY ?entity
LIMIT 200

Cross-lingual label alignment check

This federated query joins the Amharic DBpedia endpoint against the main English DBpedia endpoint to verify that Amharic entities are correctly owl:sameAs-linked to their English counterparts. Misaligned or missing links indicate entities that were extracted from Amharic Wikipedia but never connected to the global DBpedia graph.

PREFIX owl:  <http://www.w3.org/2002/07/owl#>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>

SELECT ?amEntity ?amLabel ?enEntity ?enLabel
WHERE {
  ?amEntity rdfs:label ?amLabel .
  FILTER(LANG(?amLabel) = "am")

  OPTIONAL {
    ?amEntity owl:sameAs ?enEntity .
    SERVICE <https://dbpedia.org/sparql> {
      ?enEntity rdfs:label ?enLabel .
      FILTER(LANG(?enLabel) = "en")
    }
  }
}
LIMIT 100

Property coverage audit — mapped vs. unmapped

Run against the set of 595 canonical DBpedia properties in our dataset, this query identified which properties appeared at least once in the Amharic graph and which were completely absent. The output was used to prioritise which templates to map next.

PREFIX dbo: <http://dbpedia.org/ontology/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>

SELECT ?property (COUNT(?value) AS ?usageCount)
WHERE {
  ?entity ?property ?value .
  FILTER(STRSTARTS(STR(?property), STR(dbo:)))
  FILTER EXISTS {
    ?entity rdfs:label ?label .
    FILTER(LANG(?label) = "am")
  }
}
GROUP BY ?property
ORDER BY DESC(?usageCount)

What the queries revealed

Running these four queries against the live Amharic DBpedia endpoint gave me a clear picture of the data landscape. Person entities were by far the most common type, followed by Place entities. Organisation and Work entities were underrepresented, at fewer than ten percent of the Person entity count — confirming what the mapping work had already suggested: Amharic Wikipedia's biographical coverage is strong, but its coverage of institutions, creative works, and events is limited.

On the property side, the coverage audit query showed that a large fraction of the 595 canonical DBpedia properties in the benchmark dataset had zero occurrences in the Amharic graph. Partly expected — many of those properties describe entity types that are rare in Amharic Wikipedia — but it also pointed to real gaps in mapping coverage where common entity types lacked mappings for obvious properties. The cross-lingual alignment query confirmed that owl:sameAs links to English DBpedia were present for most Person entities but missing for many Place entities, meaning those articles had been extracted but never linked back to the global graph.

Libya article and Wikipedia page refactoring

The SPARQL work showed what was missing. The Wikipedia refactoring was my attempt to close part of that gap by improving the source articles. I picked the Libya article — ሊቢያ in Amharic — as the primary test case because it's a geographically prominent country article that was using a partially incorrect infobox template. The infobox was missing the population field, had an incorrectly named government-type property, and used a manually-typed independence date the extraction framework couldn't parse because it was in Ethiopian calendar format without the right template wrapper.

Fixing these on the Libya article walked me through the full edit cycle: find the extraction failure via SPARQL, trace it to the source template or property name in the Wikipedia infobox, make the edit on Amharic Wikipedia, then verify the corrected source would produce a valid RDF triple when the extraction framework processes the next dump.

1 ሊቢያ (Libya) — primary test case; full infobox country template
2 ኢትዮጵያ (Ethiopia) — verified existing mapping, fixed date literals
3 አዲስ አበባ (Addis Ababa) — city template; added area and population fields
4 ኤርትራ (Eritrea) — country template, cross-checked borders property
5 ሶማሊያ (Somalia) — country template, verified capital and currency
6 ጂቡቲ (Djibouti) — country template, added official languages field
7 ሱዳን (Sudan) — country template, resolved duplicate property entries
8 ደቡብ ሱዳን (South Sudan) — new article; built complete infobox from scratch

I refactored eight articles in total over the week, focusing on country-level and geographic articles since they share a template structure — they all use some variant of the country infobox — so fixing one gave me a replicable pattern for the rest. It also confirmed the extraction framework mappings I'd added in previous weeks were syntactically correct: every refactored page using a mapped template produced triples when I ran it through the extraction framework locally.

The VPN access problem begins

A recurring frustration this week was the VPN requirement for the Amharic DBpedia SPARQL endpoint. Running queries from my machine was fine while connected to the institutional VPN, but it meant the website — deployed as a static GitHub Pages site — couldn't make live SPARQL calls to the endpoint without exposing credentials or requiring visitors to be on VPN. That blocked the interactive data visualization features that were part of the original website roadmap.

I documented the problem this week, but the fix — kgproxy, a lightweight SPARQL proxy deployed on AWS with an Elastic IP — didn't come until Week 14. What I got clear this week was the actual problem: the gap between "works on my machine behind VPN" and "works for any visitor to the Amharic DBpedia website."

Natnael Yohanes

Backend AI Engineer focused on ML systems, system design, distributed systems, and blockchain infrastructure.