AfirkaLLM Part 1: Africa is 0.057% of the web. Here is how we grew it.
Before you can teach a model a language, you need its text. Most of Africa's data was mislabelled, uncrawled, or never online.

Before you can teach a model a language, you need its text. Most of Africa’s data was mislabelled, uncrawled, or never online.
This is part 1 of a four-part series: 1. Where the data comes from, 2. What the experiment found, 3. Compute and terms, 4. What is missing and how to help. Parts 2 to 4 will be published over the next few days.
None of what follows in this series is possible without a corpus to train on, and that corpus almost did not exist.
The problem. Africa is home to more than 2,000 languages, yet very little of that shows up in the text AI systems learn from. Most of that text comes from the public web, and much of it from Common Crawl, a free, open archive of billions of web pages that AI builders everywhere use as their starting point. Of all the pages in that archive, just 0.057% are in an African language: fewer than 6 in every 10,000. The archive recognizes only 28 African languages, leaving languages not in the data behind.
The first surprise: much of the “missing” data was never missing. It was mislabelled. Every page in a web archive carries an automatic label saying which language it is in, and anyone building an AI system uses that label to pick out the text they want. A Hausa page labelled as Arabic or a Chichewa forum post labelled as English never gets picked, so for AI purposes it might as well not exist. To measure how often this happens, we helped build CommonLID, a test set of 373,230 lines of real web text in 109 languages and dialects, each one labelled by a person who speaks it. On that test, the biggest general-purpose AI models score about 24 points lower, on a 100-point scale, than tools built specifically to recognize languages. So the tool doing the labelling decides how much of a language you find, and a better one can recover text that is already sitting in existing archives, waiting to be found.
So we started upstream, at the web itself. We ran a community seeding campaign: 85 contributors across 19 countries curated 700 seed URLs in 34 languages. The effect rippled far beyond the seeds. Common Crawl’s late-2025 crawl collected 28.75 million pages from 15,561 domains; this led to a 38-fold expansion in African domains, adding 343,000 African-language pages and lifting Africa’s share to its highest level ever. Some languages leapt in a single crawl.

Figure 1. Some languages leapt in a single crawl. Growth in pages collected per language in the late 2025 Common Crawl run after the community seeding campaign, compared with the previous crawl.
From that enlarged crawl, our own pipeline, AfriCC, filtered an 892,000-page shortlist down to 109,062 pages confirmed genuinely African. The language of each of the pages was agreed on by two independent detectors, AfroLID and GlotLID, and is merged into a 10.2-million-document collection where nothing of unknown language or origin gets in.
The crawl is one tributary. The full harvest is much larger. Alongside AfriCC, we gathered the best existing open African text at scale, with FineWeb-2, WURA, HPLT, MADLAD-400, GlotCC, AfroLM, and Wikipedia among the main sources, into a single collection that keeps dialects distinct and tracks every license. From about 26.6 million raw documents, removing roughly 9 million too short to use and 1 million exact duplicates left 16.5 million documents: 18.6 billion tokens across 256 languages. Every source keeps its license and provenance, and dialects are kept apart rather than melted together: a Nigerian and a Ghanaian variety are not silently merged just because a label collides.
Our focus is 40 African languages. The 256-language harvest is the wider pool we draw them from, and the shape of that pool is the problem in miniature. It is steeply skewed: Afrikaans and Kiswahili alone are about half of it, and five languages hold two-thirds. At the other end, 173 languages have under a million tokens each, a tenth of one percent of the pool between them, and 86 have under ten thousand. That imbalance, a handful of higher-resource languages drowning out the rest, is exactly what the temper lever in Part 2 tries to flatten. It is also small. General-purpose models are pre-trained on trillions of tokens; our whole African harvest is under one percent of that. (Kiswahili is counted under two language codes, swh and swa; together they are 18.8% of the pool.)

Figure 2. Five languages hold two-thirds of the harvest; half the languages have under 100 thousand tokens each. The 256 languages in the complete harvest are grouped by token size: how many languages fall in each band (left) and each band’s share of all 18.6 billion tokens (right). Our training mix for the 40 focus languages is drawn from this pool.
That skew, and that size, are the starting conditions for every experiment in Part 2.
Next in the series: Part 2. What the experiment found, coming in the next few days.
Read about the project and how you can join to contribute: app.afirkallm.org/#join
Sign up and register your language: app.afirkallm.org/apply
AfirkaLLM is a project of the Data Science for Social Impact lab, University of Pretoria. We build open datasets, models, and tools for 40 African languages, and we would love your help doing it.
This work was primarily human-created. AI was used to make stylistic edits, such as changes to structure, wording, and clarity. AI was used to edit content, such as scope, information, and ideas. AI was prompted for its contributions, or AI assistance was enabled. AI-generated content was reviewed and approved. The following model(s) or application(s) were used: Claude, Gemini.