A language earns "supported" only after: (1) ≥80% UI strings native-reviewed, (2) FLORES-200 BLEU ≥45, (3) Araba bench ≥70. Never set aspirationally.
UI Bootstrap Order (static translations)
AfriWOZ
dialoguemasakhane/afri-wozBenchmark Araba conversational quality in-language; fine-tune where underperforming
Most Araba-relevant dataset — Snuggli is a conversation product
Inkuba-Instruct
instructionLelapa-AI/Inkuba-InstructMake Araba respond helpfully in-language; instruction-tune generative model
Lelapa AI (SA-based) — commercial partner candidate for Southern Africa
AfriQA
qamasakhane/afriqaAraba's comprehension + assessment/insight layer; Twi QA pairs
AfriSenti
sentimentHausaNLP/AfriSenti-twitterIn-language mood/sentiment NLU — feeds AdvancedMoodClassifier + mood insights
Direct feed into existing sentiment pipeline; grounds mood-classification in-language
FLORES-200
evaluation_benchmarkfacebook/floresGate: score each language; supported=true only when BLEU ≥45
Use to measure MT quality before shipping any language. "Gate, don't aspire."
MasakhaNER 2.0
nermasakhane/masakhaner-2In-language entity extraction — revive extractAfricanEntities with per-language NER
BibleTTS
speech_ttsmasakhane/BibleTTSFuture: voice journaling + spoken Araba in-language
CC-BY-SA ⚠️ share-alike obligation. Akuapem + Asante Twi ~86 hrs each.
WAXAL (Google Research)
speech_asr_ttsFuture: voice layer (voice journaling, spoken Araba)
~1,250 hrs transcribed speech + 235 hrs TTS across 24 Sub-Saharan languages. 2026.
Google Research Africa
WAXAL speech corpus (2026). The resource for future voice features.