gsoc-2026

GSoC 2026 Week 11

The most compute-intensive phase of the project began this week when I got H100 GPU server access via SLURM. I fully implemented Issue #3 — a configurable multi-model benchmark using Afro-XLM-R retrieval paired with local LLMs — added five new Amharic template mappings, and got the first benchmark results CSV out of the pipeline.

#gsoc-2026 #week-11 #h100 #benchmark #afro-xlmr

11 Jul 31, 2026 – Aug 7, 2026

The earlier benchmarking on Kaggle's free GPU tier had told me which models were worth investigating, but Kaggle's session time limits and notebook format made it hard to run systematic multi-seed evaluations. Week 11 changed that: my server access request to a SLURM cluster with NVIDIA H100 GPUs was approved, which meant I could run long multi-model, multi-seed evaluations and serve large local LLMs via Ollama without session interruptions. The benchmark infrastructure I built this week became the foundation for Issues #3, #4, and #5.

Getting on the H100 cluster

The server runs SLURM — the standard workload manager for HPC clusters — and GPU allocation goes through an interactive job submission. The command I used to grab a GPU for development and benchmarking sessions:

salloc --partition=capella --gpus=h100:1 --time=04:00:00

Once on a node, I started Ollama as a background process and pulled models from the Ollama registry. The key thing was that Ollama's local API — identical in interface to a remote API call — meant the DSPy-based benchmark code I'd written for Kaggle needed almost no modification to run on the cluster. The only change was pointing the LM configuration at the local Ollama endpoint instead of a public API.

Having an H100 also opened up larger models. Qwen 2.5 at 7B parameters had been the workhorse on Kaggle; the H100 let me evaluate the 72B variant, which gave noticeably better zero-shot accuracy on Amharic property mention classification.

Issue #3: the baseline benchmark

Issue #3 in the LLMIntegration repository is the first formal experiment I defined: a configurable multi-model benchmark measuring how accurately a retrieve-then-rerank pipeline can map Amharic property mentions to canonical DBpedia property labels. The dataset is dice-research/amharic-property-mapping — 2,261 training examples, 251 validation examples, and 279 test examples. The task, for each example: given an Amharic string like «ሙዚቃ ቡድን» in the context of a musician entity, select the correct DBpedia property from a set of 595 candidates.

Configurable model registry

Issue #3 introduced a YAML-based model registry so experiments can specify any combination of retriever and LLM without modifying Python source. A single config file controls which models are loaded, what shot counts to evaluate, how many random seeds to run, and where to write results.

Afro-XLM-R as the retrieval backbone

The retriever encodes all 595 canonical DBpedia property labels into a fixed index at startup. For each test example, it encodes the Amharic property mention and computes cosine similarities, returning the top-10 most similar labels as candidates for the LLM stage.

Answer snapping with RapidFuzz

LLM outputs do not always exactly match a valid property label. Answer snapping normalises the LLM response using RapidFuzz token_sort_ratio ≥ 85, falling back to the top-1 retriever result if no candidate exceeds the threshold. This made accuracy metrics stable and comparable across models.

Fuzzy accuracy metric

The evaluation metric lowercases, collapses whitespace, applies Amharic homophone normalisation via amseg, then checks for exact match or fuzz score ≥ 85 between the predicted label and the ground truth. All 279 test examples are evaluated on each seed and the results are averaged.

Models evaluated in Issue #3

I designed the benchmark to be model-agnostic — any LLM accessible via Ollama and any sentence-transformer model accessible via HuggingFace can be plugged in by changing the config file. For the baseline experiment I evaluated four models:

Qwen 2.5 (7B)

Primary LLM for few-shot classification

Strong instruction following; evaluated at 0-, 1-, 3-, 5-, and 8-shot settings.

Llama 3.1 (8B)

Secondary LLM baseline

Good zero-shot performance as a comparison anchor for the Qwen results.

Gemma 2 (9B)

Multilingual capability probe

Evaluated for its ability to handle Ge'ez script directly without translation.

Afro-XLM-R

Dense retriever (encoder)

Encodes all 595 DBpedia property labels and the query mention into 768-d vectors; cosine similarity selects the top-10 candidates.

First results CSV

By the end of the week the pipeline had run successfully across all shot settings (0, 1, 3, 5, 8) for both primary LLMs, three seeds per setting to measure variance. The output was a CSV file per model with accuracy, standard deviation, and per-example predictions. Key findings from the first run:

  • The pure retriever (Afro-XLM-R top-1) achieved approximately 52% accuracy on the 279-example test set, establishing the retrieval-only floor.
  • Adding the LLM reranker on top of the top-10 retrieval shortlist pushed accuracy to the 58–62% range at zero-shot, confirming that the two-stage design beats retrieval alone.
  • Few-shot examples (3–5 demonstrations) improved accuracy by roughly 3–5 percentage points over zero-shot, with diminishing returns beyond 5 shots.
  • Variance across seeds was low (<1 percentage point standard deviation), indicating that the shot sampling strategy was stable enough for reliable comparisons.

Five new template mappings

Alongside the benchmarking infrastructure work, I added five new Amharic template mappings to the DBpedia mappings repository, chosen from the SPARQL coverage audit in Week 10, which flagged entity types that show up frequently in Amharic Wikipedia but had incomplete or missing mapping coverage. I tested each new mapping against a real Amharic Wikipedia article using the extraction framework to confirm the properties serialised correctly into RDF triples before submitting it.

Natnael Yohanes

Backend AI Engineer focused on ML systems, system design, distributed systems, and blockchain infrastructure.