gsoc-2026
GSoC 2026 Week 4
A split week — navigating final exams while still making progress on templates, the website's data display, and the first serious look at which multilingual models could power the mapping pipeline.
4 Jun 12, 2026 – Jun 19, 2026
Week 4 landed in the middle of a demanding exam period, so I had to split time between coursework and the project. Even with less time available, it turned out to be one of the more important weeks: I had a real conversation about deployment, kept the ontology mapping work moving, redesigned the website's data display for a clearer information hierarchy, and started seriously looking into which multilingual models could power the pipeline.
Deployment discussions
I met with Thomas Tsoru, a DBpedia maintainer at DICE Research, to plan deploying the static Amharic DBpedia website. Up to this point the site had only ever run on localhost and in GitHub Actions preview builds. We talked through what it would actually take to get it on a dedicated domain — hosting requirements, who owns the DNS configuration.
A stable public URL matters beyond looks: it's what lets Amharic-speaking contributors actually find and use the site, lets external links in Wikipedia articles point somewhere real, and gives the project an identity that survives a GitHub username change. This was the first concrete step toward making the work public instead of just technically done.
Templates, wiki pages, and mappings
I built three new Amharic infobox templates, each paired with a real Amharic Wikipedia page and a DBpedia mapping entry — same pattern as Week 2: build the template, verify it on a real article, register the mapping so the extraction framework can process it. Every mapped template is a direct contribution to the Amharic DBpedia graph, since the mapping is what turns a Wikipedia infobox into structured Linked Data triples. It's unglamorous work, but it compounds: each mapping adds to the previous ones and slowly grows the share of Amharic Wikipedia content DBpedia can actually represent.
Refactoring the website literals UI
I redesigned how statistics and literal values show up on the website. Before, it was just raw numbers — a triple count, an entity count — with no explanation of what they meant. Looking at how Wikidata presents its property values gave me a clearer model: show the value with context, so it's obvious what a number refers to and why it matters.
Now every statistic is clickable and shows a tooltip explaining what the number actually means — hover over "6,412 triples" and you get "RDF triples extracted from the latest Amharic Wikipedia dump — each triple is a subject–predicate–object statement about a real-world entity." Small change to implement, but it makes a real difference for visitors who are curious about the project but don't know knowledge-graph terminology.
Exploring multilingual open-source models
My longer-term goal is a pipeline that can automatically suggest DBpedia ontology mappings for Amharic infobox fields, which means I need a model that understands Amharic text and can relate it to English-language ontology labels. This week I started searching systematically for open-source models with those capabilities.
The criteria were clear: it has to support Amharic (a low-resource language in the Ge'ez script), produce embeddings or classifications that work cross-lingually (DBpedia property labels are in English), and be small enough to run without a proprietary API. A few candidates came up — Afro-XLM-R, LaBSE, multilingual-e5, and others — that I'd formally evaluate in Week 5. This week was reading the papers, understanding how each was trained, and building a shortlist.
LangGraph agent workflow
I implemented the first end-to-end agent workflow: Afro-XLM-R as the retriever feeding into a local LLM for property mapping. It follows a retrieve-then-reason pattern — Afro-XLM-R encodes the Amharic infobox field into a dense vector, retrieves the most semantically similar DBpedia property labels from a pre-built index, and hands the shortlist to the LLM, which picks the best match and explains why.
Building this in LangGraph was worth it because its node-and-edge model makes it easy to add steps later — a translation node, a re-ranker node, a confidence-scoring node — without restructuring the whole pipeline. The first version was rough, but it connected the embedding layer to the LLM reasoning layer for the first time, which was the hardest architectural step to get past. Everything in Weeks 5–13 builds on this.