African languages barely exist in the data that AI learns from.
Africa is home to more than 2,000 languages — one of the most linguistically diverse regions on earth — yet the 28 African languages tracked by Common Crawl make up just 0.057% of its pages. Because AI models learn from data, languages that aren't in the data are left behind.
What AfriGemma does
We build the missing foundations for African-language AI: we collect, clean, and verify African web text at scale, generate instruction data grounded in African contexts, and adapt Google's Gemma models to reflect these languages and the low-resource settings they live in. We also build the evaluation benchmarks, the data infrastructure, and the review platform the work depends on.
Everything we produce — the datasets, the models, the tools, and the platform that manages them — is released openly for others to build on. Alongside the technical work we run a public seminar series on large language models that regularly draws 50+ researchers.
Growing Africa’s share of the open web
African languages are missing from the web itself, so we started upstream. With Common Crawl we ran a community seeding campaign, then a targeted extraction of our own — improving the raw material for everyone who builds on the open web, not only us.
A community seeding campaign
85 contributors across 19 countries curated 700 seed URLs covering 34 African languages. Common Crawl used these seeds in its late-2025 crawl, which collected 28.75 million pages from 15,561 domains — a 38-fold expansion in African domains — adding 343,000 African-language pages and lifting Africa's share of Common Crawl to its highest level ever.
Languages that grew the most, in a single crawl
Our targeted extraction
From that enlarged crawl we ran AfriCC, our own extraction pipeline: we sifted an 892,000-page shortlist and confirmed 109,062 genuinely African pages — each language verified by two independent detectors — then merged them with a re-verified public web dataset into the 10.2M-document collection.
From the open web to a fine-tuned model
Raised Africa's share of the web itself
A community seeding campaign with Common Crawl: 85 contributors across 19 countries curated 700 seed URLs in 34 languages. Common Crawl's late-2025 crawl collected 28.75M pages from 15,561 domains — a 38-fold expansion in African domains. Tswana grew +279%, Sango +259%, Luganda +160%. This helps everyone who builds on Common Crawl, not just us.
Built and cleaned a large African text collection
Our targeted crawl (AfriCC) recovered 109,062 genuinely African pages from an 892,000-page shortlist, and we re-verified a public web dataset (10,146,834 documents) — merged and de-duplicated into one ~10.2M-document collection.
Made sure we know what every document is
Two independent language detectors must agree on every document. We built CommonLID, a benchmark of 373,230 human-annotated web lines in 109 varieties, which shows frontier AI models trail specialised tools by ~24 F1 on African languages — evidence that much "missing" African data is really mislabelled data that careful filtering recovers.
Understood what the collection contains
No tool could sort African web pages by topic, so we trained our own (91% accurate) and ran it over all 10.2M documents. Most of the collection is everyday web writing, not news — which keeps our description of the data honest.
Audited the public instruction data
We catalogued 355 public instruction datasets and standardised 164 of them — ~367M examples across 108 languages. Mislabelled tags were the biggest distortion: text tagged "unknown" fell from 29% to 0.2% after our audit. Every fix is a re-runnable script.
Generated instruction data grounded in Africa
58,634 examples across 40 languages through six pipelines — African-entity Q&A, translation with round-trip checks, self-instruction, code-switching, cross-lingual pivoting, and functional-genre writing — with native-speaker review built in.
Started AfriGuard, a safety corpus
422,375 review items across 40 languages, stratified and classified — the first safety corpus for African languages. Fewer than 0.1% of model outputs show safe-refusal behaviour in these languages today, which is exactly why it must exist. Reviewers work under consent gates.
Chose the base model by measurement
We measured how efficiently four models' tokenizers handle our 40 languages. Gemma 4 is best (2.36 tokens/word) and the only one that keeps Amharic and Tigrinya usable — Ge'ez script costs 8+ tokens/word elsewhere versus 2.69–3.14 under Gemma 4.
Built a platform that runs the whole pipeline
Dataset ingestion, language coverage, tokenizer stats, training launches with full provenance, and native-speaker review — the unglamorous half of the work, made auditable end to end.
Mapped the African speech-data landscape
Catalogued 77 speech datasets across 394 dataset-language pairs — licenses, consent, speakers, and recorded-versus-transcribed hours — before committing compute.
The data, in numbers
Pretraining collection
Instruction & task data
By task type
| Task | Datasets | Examples |
|---|---|---|
| loading… |
Safety & speech
What the collection is about
Most African web text is everyday life, not headlines — labelled honestly instead of forced into a category.
First fine-tuned Gemma models
We settled on NVIDIA's NeMo AutoModel and fine-tuned Gemma with lightweight adapters on our own instruction data. Even the pilot — six languages, a tiny budget — delivered large gains, including +53 F1 points on Igbo question answering.
| Language | Q&A (F1) | Translation (ChrF++) | Topic acc. |
|---|---|---|---|
| Hausa | +37.1 | +2.9 | +10.5 |
| Igbo | +53.3 | +6.8 | +40.3 |
| Kinyarwanda | +31.5 | +0.4 | +28.8 |
| Swahili | +17.3 | +12.7 | +24.5 |
| Yoruba | +4.4 | +3.4 | +30.6 |
| Zulu | +45.9 | +7.7 | +21.2 |
Gains over the base Gemma model. Pilot: Gemma E2B, LoRA (rank 16), two epochs, AfriInstruct.
Scaling carefully before the big run
We are now scaling in steps — 6-, 10-, and 14-language runs under a fixed budget — to find a pipeline that holds before committing to the main pretraining pass. In parallel, we test ways to keep the model's English reasoning intact while it gains African languages, on small models first, so the lessons are cheap.
Nothing of unknown origin gets into the model.
- Every document's language is confirmed by two independent detectors that must agree.
- Each document is marked African or not; non-African text is excluded — 503 documents removed from the training set as a governance check.
- Mislabelled data is the real bottleneck: text tagged "unknown" fell from 29% to 0.2% after our audit, and every fix is a reviewable, re-runnable script.
- Native-speaker review is built into data generation — early Yoruba review caught real errors (a bird term used for insects, inconsistent tonal diacritics) that reshaped the next round.
- Reviewers work under consent gates and opt-outs; licenses and provenance are tracked for every dataset.
- The data, the tools, and the methods are openly documented and reproducible.
What comes next
- Benchmark the 14-language and per-language-budget models, and release the first fully fine-tuned Gemma as a public baseline.
- Run pretraining data-mix experiments — three mixes at 10 billion tokens each — to pick the recipe.
- Run the main pretraining pass at 100–200 billion tokens on the winning mix, then fine-tune it for the headline model.
- Extend the tokenizer's vocabulary for Amharic and Tigrinya in parallel.
- Scale native-speaker review into a trusted gold evaluation set (~400 records) and use it to sharpen automatic filtering.
- Publish the AfriCC dataset paper, and design the speech pathway toward a speech-capable model.
Support and infrastructure
Local hardware carries the fast-iteration work — pipeline development, adapter prototypes, the language-ID and topic classifiers, evaluation, and the review platform — so cloud credits go only to runs already de-risked locally. The one thing local hardware cannot do is the 100–200-billion-token pretraining run, which needs accelerator capacity only the cloud provides.
On-premise
A DGX Spark (2025 NVIDIA grant), two A6000 GPUs, and an RTX 4090 node, brought under shared Slurm scheduling with a common storage pool.
Cloud
Google Cloud compute for gathering, preparing, and training data; Vertex AI / Gemma credits for fine-tuning and evaluation across all 40 languages.
Partners we draw on
- Common Crawl — the community seeding shared task and ongoing collaboration on African web data.
- NVIDIA academic grant (2025) — access to a DGX Spark for local training and inference.
- AI4D African Languages Lab — data collection and human-capacity development.
- African Next Voices — the largest human-validated African speech collection, feeding the speech workstream.
Help us build better AI for African languages.
We are creating a high-quality dataset across 40 African languages to improve language technology and deepen linguistic research. To make sure our AI models are accurate and fluent, we need the people who know these languages best — their speakers.
What you'll do
- Verify data. Review samples to make sure they sound natural and correct.
- Guide accuracy. Confirm that automated systems align with human judgment.
- Improve representation. Make AI more effective and inclusive for your language.
- Create safe AI. Help ensure the technology is safe and culturally resonant.
Benefits for contributors
- Co-authorship. Be added as an author on the resulting dataset research paper.
- Research opportunities. Potential to coordinate further studies based on the insights gathered.
- Long-term collaboration. Explore future participatory research initiatives.