We assemble high-quality, openly-documented training data and language models for 40+ African languages that today's AI systems barely cover — with every document's language verified and its origin recorded.
Two bodies of data that work together: raw text to teach the languages, and instruction examples to make a model genuinely useful in them.
— documents of African web text, gathered by our own targeted crawl plus a re-checked public dataset, then cleaned and de-duplicated into — training tokens.
— examples across — datasets and — languages — translation, question-answering, instruction-following, conversation, and more.
We built a topic classifier (— accurate) and ran it over the whole collection, so we can describe what the data actually contains rather than guess.
The data foundation is in place; we are moving from building the data to training the models.
Most African web text is everyday writing, not news — a reality check that keeps our description of the data honest. (Topics shown for the news-covered languages.)