afirka · open African-language AI
Open data & models for African languages

Building AI that speaks Africa's languages

We assemble high-quality, openly-documented training data and language models for 40+ African languages that today's AI systems barely cover — with every document's language verified and its origin recorded.

40+
African languages targeted
documents in the cleaned collection
training units (tokens) prepared
instruction / task datasets
languages across all our data

What we have built

Two bodies of data that work together: raw text to teach the languages, and instruction examples to make a model genuinely useful in them.

Pretraining collection

documents of African web text, gathered by our own targeted crawl plus a re-checked public dataset, then cleaned and de-duplicated into training tokens.

Instruction data

examples across datasets and languages — translation, question-answering, instruction-following, conversation, and more.

Understanding the data

We built a topic classifier ( accurate) and ran it over the whole collection, so we can describe what the data actually contains rather than guess.

Progress

The data foundation is in place; we are moving from building the data to training the models.

Built on trust

  • Every document's language is confirmed by two independent detectors that must agree.
  • Each document is marked African or not; non-African text is excluded from training ( documents removed as a governance check).
  • Everything is openly documented and reproducible, and shared for the wider community.