Multilingual LLM Benchmarking
Mapping Amharic Wikipedia infobox properties to the English DBpedia ontology using a retrieve-then-rerank pipeline — five models, four experiments, Amharic script.
Slide 02 · The Problem
Mapping Amharic script to a knowledge graph.
DBpedia is a structured knowledge base extracted from Wikipedia. For Amharic Wikipedia to be part of it, every infobox property written in Amharic script must be linked to a canonical English DBpedia property. There are 595 possible targets.
የሰዎች ሀገር's የስነ ህዝብ dbo:populationTotalሀገር's ርዕሰ ከተማ dbo:capitalሙዚቀኛ's የተወለደበት ቀን dbo:birthDateሀገር's ስፋት dbo:areaTotalDataset
dice-research / amharic-property-mapping<entity type>'s <property mention>Slide 03 · System Design
Retrieve, then rerank.
Searching all 595 labels per LLM call would exceed context limits. A dense retriever narrows the field to 10 candidates in milliseconds; the LLM then picks one.
token_sort_ratio ≥ 85). Falls back to top-1 retriever result if nothing matches.Slide 04 · The Mathematics
How retrieval and evaluation work.
Cosine similarity over 595 labels
scores = premise @ labels.T
top10 = torch.topk(scores, k=10)Fuzzy-match accuracy
Normalise → lowercase + whitespace collapse + Amharic homophone normalisation (amseg),
then compare with RapidFuzz token_sort_ratio. N = 279.
Mean ± std over seeds
High σ means the result is demo-sensitive, not robustly learned.
Majority vote
Tie-break: the answer ranked highest by the retriever wins — a principled fallback because retriever rank reflects embedding similarity.
Slide 05 · Experiment 1 · Multi-Model Baseline
Five models, one retriever.
Each model evaluated at 0-shot through 5-shot with 3 random seeds per setting. Few-shot columns show mean ± std across seeds. N = 279 test examples.
| Model | 0-shot | 1-shot | 3-shot | 5-shot |
|---|---|---|---|---|
| Qwen 2.5 32B | 57.35 | 59.02 ±1.8 | 59.85 ±1.5 | 58.18 ±1.9 |
| Gemma 2 9B | 57.71 | 60.22 ±2.1 | 61.65 ±2.4 | 65.83 ±2.84 |
| Llama 3.1 8B | 42.65 | 50.9 ±3.2 | 61.41 ±2.1 | 56.87 ±2.9 |
| Llama 3.2 3B | 40.5 | 52.69 ±4.1 | 47.19 ±4.2 | 52.09 ±4.3 |
| Aya Expanse 8B | 37.28 | 48.99 ±5.2 | 48.99 ±5.8 | 36.8 ±8.1 |
Cyan = best per model. Aya Expanse 5-shot (36.80 ± 8.1) is worse than 0-shot.
Slide 06 · Experiment 2 · Prompt Engineering
Four ways to ask the LLM to choose.
Same retriever shortlist across all strategies. Only the prompt structure changes. Best result: Gemma 2 + chain-of-thought at 5-shot → 67.62%.
Slide 07 · Prompt Engineering Results
CoT at 5-shot reaches 67.62%.
Gemma 2 is the only model that consistently improves with every strategy.
| Model | Strategy | Shots | Accuracy | Std |
|---|---|---|---|---|
| Gemma 2 | cot | 5 | 67.62% | ±0.89 |
| Gemma 2 | persona | 5 | 66.19% | ±3.47 |
| Gemma 2 | translation-cot | 3 | 64.40% | ±2.65 |
| Llama 3.1 | persona | 3 | 63.80% | ±0.78 |
| Qwen 2.5 32B | direct | 2 | 60.45% | ±5.51 |
| Llama 3.2 | cot | 4 | 56.87% | ±2.16 |
| Aya Expanse | direct | 1 | 48.27% | ±7.82 |
Gemma 2 · CoT grid
Slide 08 · Experiment 3 · Cross-Lingual Translation
Should the LLM see Amharic or English?
The same model translates its own input, then reranks. Translation runs once up front and is reused across all technique/shot/seed combinations.
Pivot
The Amharic premise is fully replaced by its English translation. The LLM never sees Amharic script.
"country's population"[dbo:populationTotal, …]Tests a pure English pipeline. Translation errors propagate directly — no Amharic signal to fall back on.
Augmented
The Amharic premise is kept and the English translation is added as a separate reference field alongside it.
ሀገር's ህዝብ"country's population"Richer signal — the model can use whichever representation it trusts more. Augmented beats Pivot on every model at every shot count.
Slide 09 · Translation Results
Augmented wins — context beats replacement.
Pivot (English only)
| Model | 0-shot | Best |
|---|---|---|
| Qwen 2.5 32B | 53.76 | 52.93 (5s) |
| Gemma 2 | 55.91 | 61.53 (4s) |
| Llama 3.1 | 36.56 | 55.20 (3s) |
| Llama 3.2 | 41.94 | 53.17 (5s) |
| Aya Expanse | 39.07 | 42.17 (3s) |
Augmented (Amharic + English)
| Model | 0-shot | Best |
|---|---|---|
| Qwen 2.5 32B | 57.35 | 54.12 (3s) |
| Gemma 2 | 50.54 | 66.90 (4s) |
| Llama 3.1 | 53.41 ↑ | 59.02 (4s) |
| Llama 3.2 | 47.67 ↑ | 56.99 (4s) |
| Aya Expanse | 42.29 ↑ | 53.65 (1s) |
Key findings
- Helps weaker models. Llama 3.1 0-shot: 42.65% → 53.41% (+10.76 pp). Models that struggle with Amharic script gain a meaningful boost from the English hint.
- Hurts strong models at few-shot. Qwen 2.5 32B drops from 60.21% → 54.12% at 3-shot. The extra field adds noise for models already capable of cross-lingual reasoning.
- Gemma 2 augmented peak: 66.90% at 4-shot. Second-highest single-model result in the project, trailing only chain-of-thought prompting.
- Augmented always beats Pivot. Keeping the Amharic context alongside the translation is better than discarding it in every case tested.
Slide 10 · Experiment 4 · Ensemble Methods
Five models vote better than one chooses.
Member predictions collected once per (shots, seed) and reused for both strategies. No recomputation on strategy switch.
Example — all 5 models predict for one test item:
| Strategy | 0-shot | 2-shot | 3-shot | 5-shot |
|---|---|---|---|---|
| vote | 63.08 | 66.19 | 66.67 | 66.19 |
| rerank-vote | 57.71 | 58.78 | 58.78 | 58.78 |
Best solo model at 0-shot: Gemma 2 at 57.71%
Slide 11 · Results Summary
Every experiment raises accuracy.
| Experiment | Best Config | Accuracy | Std |
|---|---|---|---|
| 1 · Baseline | Gemma 2 · 5-shot | 65.83% | ±2.84 |
| 3 · Translation | Gemma 2 + Augmented · 4-shot | 66.90% | ±1.50 |
| 4 · Ensemble | Vote · 3-shot | 66.67% | ±2.03 |
| 2 · Prompt Engineering | Gemma 2 + CoT · 5-shot | 67.62% | ±0.89 |
| 4 · Ensemble (0-shot) | Vote · no demos | 63.08% | — |
What we learned
- Chain-of-thought + 5 shots is the most accurate single-model strategy at 67.62% ± 0.89.
- Ensemble vote at 0-shot (63.08%) beats most individual few-shot results — useful when you have no labeled examples.
- Translation helps weaker models, hurts stronger ones. Keep Amharic context alongside the English translation.
- CoT at 0-shot hurts. Add at least 2 demonstrations before enabling chain-of-thought.
- Next step: fine-tune Afro-XLM-R on the training split — the retriever sets the ceiling for everything above it.