gsoc-2026

GSoC 2026 Week 7

The pipeline matured a lot this week. I added a top-10 dense retrieval layer before the LLM, turning it into a proper retrieve-then-rerank system, and studied LLM consistency metrics to understand how much to trust the results.

#gsoc-2026 #week-7 #rag #retrieval #reliability #few-shot

7 Jul 3, 2026 – Jul 10, 2026

Week 7 turned the pipeline from a prototype into a system with a principled architecture. Adding dense retrieval before the LLM reasoning step is the defining architectural decision so far: it constrains the LLM's search space to a manageable shortlist, which improves both accuracy and consistency by a lot. I also built a reliability evaluation framework this week, because knowing a model's average accuracy isn't enough — knowing how much that accuracy varies across runs is what actually makes it trustworthy.

LLM reliability evaluation metrics

Before I could trust the Week 6 results, I had to answer a basic question: do LLMs give the same answer when asked the same question twice? No — LLM outputs are stochastic, so accuracy figures from a single run are noisy estimates. This week I built a systematic reliability evaluation framework around three core metrics.

Self-consistency is the simplest one: run the same prompt N times on the same input and measure what fraction of runs agree. A model that returns the same answer 9 out of 10 times is far more reliable than one that agrees with itself 6 out of 10, even at similar mean accuracy. Variance across random seeds captures a related but different concern — whether accuracy changes meaningfully when the seed changes. And standard deviation of accuracy across runs gives a single number for how much to trust a reported mean: a model at 60% accuracy with ±1 pp standard deviation is practically more valuable than one at 62% with ±8 pp, since the latter's real performance could plausibly be anywhere between 54% and 70%.

Retrieve-then-rerank pipeline

The core architectural change this week was adding a top-10 retrieval step using Afro-XLM-R dense embeddings before every LLM call. The pipeline now works in two stages. First, I pre-encode all 595 DBpedia property labels into a matrix of dense vectors using Afro-XLM-R as a bi-encoder. Then, for each Amharic infobox field at inference time, I encode the field with the same model and use cosine similarity to select the 10 most semantically similar DBpedia property labels from the full vocabulary.

The LLM then only sees these 10 candidates instead of all 595. This retrieve-then-rerank architecture helps in two ways. Accuracy improves because the LLM isn't reasoning over an enormous vocabulary anymore — it just has to pick the best match from a pre-filtered shortlist where the correct answer is usually present. And consistency improves because the LLM's decision space is constrained, which cuts down on the model hallucinating a property label that isn't in the vocabulary at all.

Few-shot prompting with retrieved candidates

The retrieved top-10 candidates did double duty this week: they're both the shortlist for the LLM to choose from and the demonstration pool for few-shot examples. Instead of drawing few-shot examples randomly from the training set, I select examples whose ground-truth labels are similar to the current query's top-10 retrieved candidates, so the demonstrations are relevant to the specific disambiguation problem at hand.

The Kaggle results showed meaningful accuracy gains with 1 to 3 demonstrations over zero-shot. Moving from zero-shot to one example gave the largest single gain, with diminishing returns after that. This confirmed retrieve-then-rerank as the right foundation and settled the few-shot demonstration strategy I'd carry into the formal benchmark in Week 11.

Results documented

I ran the full pipeline — Afro-XLM-R retriever, top-10 shortlist, few-shot LLM — on the Kaggle server and documented all results systematically: accuracy per model and per shot count (0 through 5), each configuration repeated with multiple random seeds to compute the standard deviation. This documentation practice became the template for every experiment after it — model, prompt strategy, shot count, mean accuracy, standard deviation, every time. Without that structure, comparing results across weeks would've been impossible.

This week's results became the baseline for the multi-model benchmark in Issue #3, which would run on the H100 server in Week 11. Every experiment after this was designed to either beat this baseline or explain why it couldn't.

Natnael Yohanes

Backend AI Engineer focused on ML systems, system design, distributed systems, and blockchain infrastructure.