gsoc-2026
GSoC 2026 Week 5
A research-heavy week. I got the website's first major PR merged, evaluated six candidate embedding models, and worked through the prompt engineering literature systematically — landing on DSPy as a practical framework.
5 Jun 19, 2026 – Jun 26, 2026
Week 5 was the most research-heavy week so far. There wasn't one dominant task — I had to move three things forward at once: get the website's first major PR through, evaluate the model shortlist rigorously enough to narrow it to serious candidates, and read the prompt engineering literature closely enough to have a real opinion on which techniques would matter most for property mapping. By the end of the week DSPy had become the clear practical direction.
Website PR #1
I refactored the Amharic DBpedia website and submitted the changes as PR #1. The refactor broke the monolithic page down into smaller, reusable Svelte components and fixed the data presentation so the statistics and entity counts from Week 4 displayed correctly across screen sizes. Submitting a real PR instead of pushing straight to main also set the review workflow every later website contribution would follow.
Six models evaluated for Amharic property mapping
I expanded the Week 4 shortlist and evaluated it more carefully, checking each model against three criteria: documented Amharic support, cross-lingual retrieval capability, and whether it's practical to run without API access. Here are the six candidates and the reasoning behind each:
LaBSE
Google's Language-Agnostic BERT Sentence Embedding model supports 109 languages including Amharic. Unlike standard masked LM pretraining, LaBSE is trained on bilingual sentence pairs, teaching the model to produce similar vector representations for sentences with the same meaning in different languages. This cross-lingual transfer is precisely what is needed for matching Amharic infobox fields to English DBpedia property labels.
multilingual-e5-large
Built on XLM-RoBERTa-large, multilingual-e5-large uses a shared cross-lingual vector space that aligns representations from different languages into the same mathematical space. The result is that semantically equivalent phrases in Amharic and English land near each other in the embedding space, giving the model genuine multilingual semantic understanding rather than simple token-level overlap.
bge-m3
The M3 in bge-m3 stands for Multi-Lingual, Multi-Functionality, Multi-Granularity. Developed by BAAI, it supports over 100 languages including Amharic and is trained on large cross-lingual corpora covering many domains. The multi-granularity design means it can produce embeddings at the token, sentence, or passage level, giving flexibility in how retrieval is structured.
am-roberta
A monolingual Amharic RoBERTa model trained specifically on Amharic text, making it excellent at understanding the morphological richness and token structure of the Ge'ez script. The limitation is significant for this project: being monolingual, it cannot directly bridge Amharic inputs to English DBpedia property labels without an additional cross-lingual step. It remains valuable as a reference model for Amharic-only understanding tasks.
jina-embeddings-v3
Jina AI's embedding model supports 89 languages including Amharic and uses LoRA adapters under the hood for GPU-efficient fine-tuning and inference. The LoRA design means the base model can be adapted to specific retrieval tasks with a fraction of the usual compute, which is an attractive property for a project that may eventually want task-specific fine-tuning.
gte-multilingual-base
At the time of evaluation, gte-multilingual-base sat at the top of the HuggingFace MTEB multilingual retrieval leaderboard. Developed by Alibaba NLP, it is trained using contrastive learning — a technique that directly optimises the model to pull semantically similar sentence pairs together and push dissimilar pairs apart in the vector space, which is ideal for retrieval-based property mapping.
Prompt engineering deep-dive
I worked through the prompt engineering literature systematically and came out with a working taxonomy. Starting from the anatomy of a well-formed prompt — role, context, instruction, examples, output format — I covered zero-shot (instruction only), one-shot, and few-shot prompting, and why adding examples generally helps LLMs give more consistent answers on narrow classification tasks.
The more advanced techniques mattered just as much. Chain-of-thought (CoT) asks the model to reason step-by-step before answering, which helps on tasks needing multi-step inference — like explaining why an Amharic phrase maps to a specific DBpedia property. Tree-of-thoughts (ToT) extends that by exploring several reasoning branches at once. Chain-of-drafts is a newer, cheaper variant that produces short intermediate drafts instead of full reasoning chains. And ReAct — Think, Act, Observe — combines reasoning with tool calls, which maps directly onto an agent workflow where the LLM can call a retrieval tool before deciding.
GEPA paper
I read the abstract and introduction of the GEPA paper (Generative Evolutionary Prompt Adaptation). The core idea: instead of fine-tuning model weights — expensive, needs labelled data and GPU time — you can build a system that automatically rewrites and optimizes the prompt itself. It treats the prompt like a candidate solution in an optimization loop: generate variants, score them against a held-out metric, keep the best, repeat. Directly relevant here, since hand-tuning prompts is brittle and slow.
DSPy
The most important discovery of the week was DSPy, Stanford's library for programming with language models. The model is elegant: instead of hand-crafting prompt text, you write a Python program describing the logical structure of what you want the LLM to do ("given an Amharic mention and a list of candidate DBpedia properties, choose the best one"), give it a scoring metric (exact-match accuracy on a labelled dataset), and DSPy iterates, mutates, and optimizes the exact prompt wording for you — programmatic prompt engineering at scale. For property mapping, that meant I could find good prompts without guessing.
MCP for dbpedia-mapper
I implemented an MCP (Model Context Protocol) server for the DBpedia mapper, exposing two initial tool functions that make it callable from LLM agents. MCP is an open standard for connecting LLM agents to external tools and data sources, so wiring it up here means the property matching logic can be invoked by any MCP-compatible agent framework — including the LangGraph workflow from Week 4. The two tools covered property retrieval and mapping suggestion, keeping the retrieval concern separate from the reasoning concern.