ArXiv

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

Authors
Paul Azunre
Categories
cs.CL
arXiv
https://arxiv.org/abs/2607.21540v1
PDF
https://arxiv.org/pdf/2607.21540v1

Brief

DONDO presents 21 monolingual and 5 multilingual w2v-BERT 2.0 ASR base models for 27 African language varieties from six countries, fine-tuned primarily on read religious-text speech. A two-step (one family three-step) learning-rate-annealed adaptation yields average WERs of 10–13% for multilingual checkpoints, uses a one-hot prefix language-conditioning to steer a single model, and is released under Apache-2.0 on Hugging Face.

Why it matters

DONDO is a family of open ASR base models (w2v-BERT 2.0) comprising 21 monolingual and 5 multilingual checkpoints covering 27 language varieties across Ghana, Sierra Leone, Nigeria, Senegal, Kenya and Zimbabwe (paper posted 2026-07-23).

Key details

  • Models are fine-tuned mainly on read speech from religious texts using a two-step (and one three-step) learning-rate-annealed procedure that first adapts a shared multilingual model at high LR then anneals; the five multilingual families reach average WERs of 10–13%, closing much of the gap to monolingual baselines.
  • A lightweight language-conditioning mechanism prepends a one-hot language identity as prefix frames to acoustic features so a single multilingual checkpoint can be steered to target languages; all models are released on Hugging Face KhayaAI under the Apache-2.0 license and cover an estimated ~100 million first-language speakers.
Source evidence

Abstract

We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models and five multilingual models spanning twenty-seven language varieties across Ghana, Sierra Leone, Nigeria, Senegal, Kenya and Zimbabwe. Models are fine-tuned primarily on read speech drawn from religious texts, which offer broad, license-clear and orthographically consistent coverage for languages that otherwise lack transcribed audio. We describe a two-step (and, for one family, three-step) learning-rate-annealed fine-tuning procedure that first adapts a shared multilingual model at a high learning rate and then anneals it to recover, and in several cases surpass, strong monolingual baselines. We further describe a lightweight language-conditioning mechanism that injects a one-hot language identity as a sequence of prefix frames prepended to the acoustic features, allowing a single multilingual checkpoint to be steered to a target language at inference. Across the five multilingual families the annealed models reach average word error rates (WER) of 10-13%, closing most of the gap to monolingual models while covering many languages in a single checkpoint. All models are released on the Hugging Face KhayaAI organisation under the Apache-2.0 license (attribution only) so that others may fine-tune them freely, including for commercial use. We provide a conservative estimate that the languages covered are spoken by on the order of one hundred million first-language speakers, and by substantially more when second-language use is included.