gsoc-2026

GSoC 2026 Week 6

The first real benchmarking results arrived this week. I ran models on Kaggle GPU, tested prompt strategies side-by-side, and built a specialised Amharic-aware fuzzy search to handle the language's phonetic variation.

#gsoc-2026 #week-6 #benchmarking #kaggle #prompt-engineering #fuzzy-search

6 Jun 26, 2026 – Jul 3, 2026

Week 6 produced the first concrete numbers. Instead of reading about models and techniques, this week was about running them and measuring. Kaggle's free GPU tier let me run the first benchmarking suite without dedicated compute, I compared multiple prompt strategies head-to-head on a single model, and I fixed a real problem with exact string matching in Amharic by adding a language-aware fuzzy search step.

Reading the prompt engineering papers

I went back to the primary papers to ground the prompt engineering survey from Week 5. Reading the original chain-of-thought paper, the tree-of-thoughts paper, and the few-shot learning work made the trade-offs far clearer than any blog summary could. The key insight was the relationship between task complexity and the benefit of reasoning steps: simple retrieval tasks get little from CoT, but tasks requiring disambiguation between similar properties (e.g. dbo:country vs. dbo:nationality vs. dbo:birthPlace) benefit a lot from having the model articulate why it's choosing one property over another.

I also gave chain-of-drafts a serious look. The original paper's claim — that producing short intermediate drafts instead of full reasoning chains gives most of the accuracy benefit at a fraction of the token cost — is directly relevant to a production pipeline where inference latency and token budget both matter.

Deep dive into the four multilingual models

I studied the four most promising models from the Week 5 shortlist — LaBSE, multilingual-e5-large, bge-m3, and gte-multilingual-base — in technical depth, focusing on exactly how each handles cross-lingual alignment. LaBSE uses bilingual sentence pairs as direct training signal, so the alignment is explicitly supervised. The XLM-RoBERTa-based models (multilingual-e5-large, bge-m3) learn a shared representation space through masked language modelling across many languages at once, which gives broader coverage but potentially less precise cross-lingual transfer. gte-multilingual-base uses contrastive learning at the sentence level, which optimises more directly for retrieval quality than masked LM does.

For Amharic specifically, the Ge'ez script is a real challenge for models trained mostly on Latin-script languages: tokenisation, character coverage, and morphological richness all behave differently. Checking each model's tokeniser and its Amharic vocabulary size became part of my evaluation criteria.

Kaggle GPU benchmarking

I ran the full multilingual LLM benchmarking suite for the first time on a Kaggle GPU, captured in the multi-llm-benchmarking-v3 notebook. I built the pipeline as a structured Kaggle notebook that loads the labelled Amharic property mapping dataset, encodes the DBpedia property labels with the chosen retriever, and for each test example retrieves the top-10 candidates and passes them to an LLM with varying prompts. I tested multiple prompt strategies — zero-shot, one-shot, few-shot — on a single model in one run and recorded the accuracy for each combination.

Kaggle's environment has real constraints — session time limits, memory caps, no persistent background processes — that ruled out some experiment configurations. But for a first round of benchmarking it was exactly enough. The results were the first concrete evidence of which prompt strategies actually work, moving me from informed speculation to measured fact.

Server access request

Kaggle's free tier wasn't going to be enough to run the full benchmark suite — multiple models across multiple prompt strategies, with multiple random seeds for statistical reliability. I submitted an access request for a dedicated GPU server through Mentor Tilahun, asking for a minimum of 48 GB VRAM. The justification was straightforward: the largest model on the candidate list (Qwen 2.5 32B) needs about 20 GB in float16, and running five models at six shot settings with three seeds each is a combinatorial workload that needs sustained GPU compute rather than 12-hour session windows.

The request would eventually unlock H100 access in Week 11, but this week was about preparing the formal justification the cluster administrators would need before granting access.

Fuzzy search with Amharic normalisation

Exact string matching turned out to be inadequate for Amharic evaluation, so I added fuzzy search with a custom Amharic normalisation pre-processing step. The problem is fundamental to the Ge'ez writing system: many phonetically identical sounds are written with different characters across the seven vowel forms (ሀ, ሃ, ሄ, ህ, ሆ, and others all represent related sounds), and different writers or Wikipedia editors use these interchangeably for the same word.

I used amseg, a Python library for Amharic text normalisation that handles Ge'ez homophone normalisation as a pre-processing step. Before any string comparison, both the predicted label and the ground-truth label pass through amseg's normaliser, which collapses phonetically equivalent character variants to a canonical form. This one change improved evaluation accuracy noticeably, since a model giving a correct-sounding answer no longer got penalised for a different but phonetically identical spelling. Without normalisation, exact string matching missed a large fraction of genuinely correct answers.

Natnael Yohanes

Backend AI Engineer focused on ML systems, system design, distributed systems, and blockchain infrastructure.