AfirkaLLM Part 2: What teaches an AI to speak African languages?
Five controlled experiments, each with one change, judged against the untrained model. The answer wasn't what we expected.

Five controlled experiments, each with one change, judged against the untrained model. The answer wasn’t what we expected.
This is part 2 of a four-part series: 1. Where the data comes from, 2. What the experiment found, 3. Compute and terms, 4. What is missing and how to help. Parts 3 and 4 will be published over the next few days.
We are building open AI for 40 African languages by taking Google’s Gemma model and continuing its training on African-language text drawn from the harvest described in Part 1. The hardest question in that work is not whether it can be done. It is what data does the teaching. Three obvious options:
- Flatten the language proportions, so smaller languages are not drowned out? [1]
- Add a little code or maths to sharpen reasoning? [2]
- Cap the religious text that dominates many African web corpora to make room for health, agriculture, and news? [3]
Each of those has worked for other teams, so each is a reasonable guess for us too, and a wrong guess wastes weeks of expensive compute. So instead of guessing, we ran a controlled experiment: change one thing at a time and measure what happens.
We chose Gemma for a measurable reason. A model can only serve a language as well as its tokenizer represents it. A tokenizer chops text into the small pieces, called tokens, that a model actually reads; the fewer pieces a word needs, the cheaper that language is to learn. Gemma’s tokenizer is the most efficient of the leading options across our 40 languages (2.36 tokens per word), and it keeps Ge’ez-script languages workable. Amharic and Tigrinya cost about 3 tokens per word under Gemma versus 8 or more elsewhere, the difference between a language a model can afford to learn and one it cannot.

Figure 1. Gemma’s tokenizer is the cheapest for our languages. Tokens per word by tokenizer, averaged over the 40 focus languages and for the two Ge’ez-script languages; lower is better. Qwen 3 was not measured on Ge’ez-script languages.
The experiment
We took Gemma-4-E4B and continued its pre-training five times on a training mix built around our 40 focus languages. The mix also carries 16 more languages: dialects and closely related varieties of the focus languages and anchor languages that help the model learn them. Each run saw exactly 5 billion tokens, on identical settings, and each changed one thing against a shared baseline. Researchers call runs like these ablations.
| Run | The one change | The question it answers |
|---|---|---|
| Baseline | The reference mix (uniform sampling) | What does “no change” look like? |
| Temper | Flatten language proportions | Do bigger African languages stop starving the smaller ones? |
| Code | Add a code anchor | Does code make the model reason better, and at what cost? |
| Math | Add a maths anchor (FineMath) | Does maths data make the model better with numbers? |
| Religious cap | Limit religious text per language | Does a wider mix of topics beat a religion-heavy one? |
Because only one thing differs per run, any difference in the results is attributable to that lever. Five clean experiments instead of a hunch.
Reading the results honestly matters as much as running them. A model’s score on its own test data cannot be compared with another model’s, so we never use it for ranking. Every run is judged on the same fixed external benchmarks: AfriMMLU and AfriMCQA (knowledge), AfriXNLI (inference), Belebele (reading comprehension), AfriMGSM (grade-school maths), and FLORES and AfriScience MT (translation), plus an English “guardrail” that checks the model did not forget English while learning African languages. Most are multiple-choice tests scored by accuracy; the two translation tests use ChrF, a 0 to 100 score of how closely a translation matches a human one. Where a benchmark scores close to random guessing, we say so rather than pretend the differences mean something.
What we found
The most important number in this experiment turned out to be the one that is easy to forget: the “before.” We evaluated the untrained base model on exactly the same benchmarks, and only against that reference does the result become honest.
Continued pre-training, at this budget, trades knowledge for translation.

Figure 2. Continued pre-training buys translation and spends knowledge. Change in score after 5 billion tokens against the untrained Gemma-4-E4B base, for every benchmark and run; copper is worse, blue is better. Accuracy rows in percentage points, translation rows in ChrF points.
- On translation (FLORES and the scientific translation dataset, AfriScience MT), continued pre-training clearly helps. All five runs beat the untrained base. This is the real, measurable win: more African text in, better African translation out.
- On knowledge multiple-choice (AfriMMLU), the untrained base beats every ablation. Base AfriMMLU is 0.38; the ablations land between 0.28 and 0.36. In other words, 5 billion tokens of raw African web text erodes what the model already knew. The model forgets general skills while it fits a narrower slice of text.
That trade-off is the finding. It would have been invisible without the before-column, and reporting it, rather than only the best ablation, is the difference between marketing and science. The full scores behind the matrix:
| Benchmark | Untrained base | Baseline | Temper | Code | Religious cap | Math |
|---|---|---|---|---|---|---|
| AfriMMLU (accuracy) | 0.381 | 0.318 | 0.305 | 0.282 | 0.352 | 0.357 |
| AfriMCQA (accuracy) | 0.277 | 0.276 | 0.271 | 0.268 | 0.274 | 0.258 |
| AfriXNLI (accuracy) | 0.359 | 0.359 | 0.368 | 0.354 | 0.349 | 0.372 |
| Belebele (accuracy) | 0.302 | 0.290 | 0.286 | 0.283 | 0.295 | 0.292 |
| AfriMGSM (accuracy) | 0.018 | 0.024 | 0.030 | 0.018 | 0.024 | 0.021 |
| FLORES (ChrF) | 26.66 | 29.61 | 28.81 | 26.79 | 29.14 | 29.13 |
| AfriScience MT (ChrF) | 27.35 | 33.54 | 32.06 | 27.49 | 30.46 | 30.33 |
Which lever forgets the least? That is what the ablations are really choosing between, and the honest answer is that the differences between the mixes are small. So we measured every one against the plain baseline (uniform mix, no lever), not just against each other.
- The plain baseline is the strongest control, not a punching bag. It is the best translator of the six on both translation benchmarks (FLORES 29.6 ChrF, AfriScience MT 33.5 ChrF) and the best at holding onto English: its English MMLU even edges the pre-training base. Any lever has to beat this, not the untrained model, and once you look, almost none of them clearly do.
- Capping religious text is the one lever with a real, if modest, case. It edges the baseline on knowledge (AfriMMLU 0.352 against 0.318) while staying close on translation. Its whole value is losing slightly less knowledge. It does not translate better than the plain baseline.
- Both anchors failed the specific thing they were added to do. The maths anchor was included to lift numeric reasoning, but on AfriMGSM it scores 0.021, no better than the plain baseline’s 0.024. Both are effectively zero: no model does African grade-school maths at this budget. It is also the worst run on native-language multiple choice (AfriMCQA 0.258). Its FineMath data is English, so it bought lower English perplexity (a measure of how confused the model is by ordinary English; lower is better) and nothing African. The code anchor is worse still: lowest on nearly every African task and the furthest English drift, with English perplexity blown out to 58.6 against roughly 25 to 30 for the others. Both anchors spend African tokens and buy nothing for African languages.
- Tempering the language mix does not beat the plain baseline. Against it, temper is lower on knowledge and translation, and it drifts on English. An earlier read that called it a “cheap win” was comparing it only to the other levers, with no baseline in the picture. Once the baseline is there, the win disappears.

Figure 3. The plain baseline kept its English; the code mix lost the most. English perplexity (lower is better) and English MMLU accuracy (higher is better) after 5 billion tokens of African text. The line marks the baseline on the left and the untrained base on the right. The math run’s low perplexity reflects its English FineMath anchor, not African gain.
A caveat we hold to: at this eval size, several benchmarks (AfriXNLI, Belebele, AfriMGSM, AfriMCQA) score close to random guessing for every run, so we do not rank the levers on them. What is not noise, and what this blog leads with: the untrained base beats all on knowledge (AfriMMLU); all five ablations beat the base on translation (FLORES and AfriScience MT); the plain baseline is the best translator and English retainer; and both the code and maths anchors fail at the one thing they were added for.
The translation gain is not spread evenly across languages. The two languages that dominate the corpus, Afrikaans and Kiswahili, barely move: they were already strong, and adding more of them taught the model nothing new. The gains go to the languages with less data. Luganda climbs 21 ChrF points under the baseline, Kabyle 19, Kikuyu 16 under temper, and Oromo and Lingala each about 20 under the religious cap. But the same matrix shows the cost of a 5-billion-token run on a skewed corpus: Shona loses 20 points under the religious cap, and Chichewa 15 under the baseline and code runs. An average across languages hides both.

Figure 4. Translation gains land in the low-resource languages, and not evenly. Change in FLORES ChrF against the untrained base, per language and run, English into each language, ordered by base score (right column). Copper is worse than the base; blue is better.
A subtle but decisive detail: across the runs, lower training loss did not mean a better model. Loss is the score a model tries to drive down during training. The code mix reached the lowest loss (code is easy to predict) yet was the worst on the benchmarks. The more varied religious-capped mix ran at the highest loss yet scored the best of the levers. So training loss cannot pick the recipe. And because every run’s loss had already plateaued by 5 billion tokens, the loss of knowledge is about what the model trained on, not how long. More tokens of the same data will not recover it.

Figure 5. Lower training loss did not mean a better model. Each run’s final training loss against its AfriMMLU score after 5 billion tokens; the easiest text to predict produced the worst model. Every loss curve had flattened out by then.
Where this goes next
The experiment did its job. It narrows the recipe to a clear verdict: religious cap in, both anchors (code and maths) out, tempering unproven. But the honest headline is that continued pre-training alone, at 5 billion tokens, buys translation at some knowledge cost, and no lever recovers the base’s knowledge on its own. Recovering that is the job of the next stages: a gentler, longer full run with measures against forgetting (a wider mix of topics, a slice of English kept in the mix, and freezing the model’s untouched vision and speech parts), and the instruction tuning that follows, which re-teaches the model to answer rather than just speak. The comparison that ultimately matters, the combined model after fine-tuning against the base, is the one still to come.
Further reading
[1] Training on many languages at once lets the biggest ones crowd out the rest, so multilingual models routinely “flatten” the mix by sampling small languages more often than their raw share. Google’s mT5 used the same setting we tested (α = 0.3): Xue et al., 2021, arXiv:2010.11934. See also Conneau et al., 2020 (XLM-R), arXiv:1911.02116, and Arivazhagan et al., 2019, arXiv:1907.05019.
[2] Several studies find that adding computer code or maths to a model’s training text improves its general reasoning, not just its coding or arithmetic: Aryabumi et al., 2024, “To Code, or Not To Code?”, arXiv:2408.10914, and Shao et al., 2024, DeepSeekMath, arXiv:2402.03300.
[3] An audit of web-crawled multilingual datasets found that much of the text for low-resource languages is mislabelled, low quality, or religious: Kreutzer et al., 2022, “Quality at a Glance,” arXiv:2103.12028. Bible and Jehovah’s Witness translations are among the largest freely available sources for many African languages, for example, JW300: Agić and Vulić, 2019, ACL Anthology P19-1310.
Next in the series: Part 3. Compute and terms, coming in the next few days.
Read about the project and how you can join to contribute: app.afirkallm.org/#join
Sign up and register your language: app.afirkallm.org/apply
🎥 Watch: LLM Building and Sovereign AI. In 2026 we have also been running a special theme of our Data Science for Society Seminar Series focused on LLM building and sovereign AI. Recordings are available on YouTube: watch the LLM Building and Sovereign AI playlist, and subscribe to the channel so you don’t miss new sessions.
AfirkaLLM is a project of the Data Science for Social Impact lab, University of Pretoria. We build open datasets, models, and tools for 40 African languages, and we would love your help doing it.
This work was primarily human-created. AI was used to make stylistic edits, such as changes to structure, wording, and clarity. AI was used to edit content, such as scope, information, and ideas. AI was prompted for its contributions, or AI assistance was enabled. AI-generated content was reviewed and approved. The following model(s) or application(s) were used: Claude, Gemini.