Millions of people speak Lingala and Shona fluently without ever having learned to read or write them. That makes voice the obvious way for them to use technology, and the hardest one to build, because most speech-recognition systems cannot understand either language.
A challenge run by Google Research and the data-science platform Zindi set out to change that, and Google published the results on 22 September 2026. It drew 1,462 entrants from 100 countries, 42 of them African, who submitted 8,837 solutions and, by Google’s count, put in more than 24,000 hours of work.
Africans take second and third place
First place went to Team Pasketti, from Kazakhstan and Canada, with an ensemble of 8 models that check each other’s transcriptions. Second place went to Alban Nyantudre, a machine-learning engineer from Burkina Faso who builds tools for the Moore language, and third to Abdourahamane Ide Salifou, an artificial intelligence engineering student from Niger. Google said the top 3 scores were close, each reached by a different method.
“Plenty of people I grew up around speak their language perfectly but cannot read or write it,” Nyantudre said. Google did not disclose the prize amounts.
The dataset behind it
Entrants built their systems on WAXAL, an open collection of African-language speech that Google released on 2 February 2026 after 3 years of development. Google’s own descriptions of its size vary: 21 languages at launch, 27 in a Google Research post in March, and 32 in this month’s results.
According to the March post, the data was gathered with Makerere University, the University of Ghana, Digital Umuganda, Addis Ababa University, the African Institute for Mathematical Sciences in Senegal, Media Trust, and Loud n Clear, and covers languages spoken by more than 100 million people.
The data is published on Hugging Face under Creative Commons licences, which allow commercial reuse with attribution, and its dataset card credits funding from Google and the Gates Foundation. That openness is the point: a startup in Kinshasa or Harare can build on it without negotiating access.
Part of a wider push
WAXAL is one of several efforts to give African languages the training data that English, French and Portuguese take for granted. In May, the GSMA and Pleias released CommonLingua, covering 61 African languages, and Zindi has run an AI safety challenge with the GSMA to test how language models behave for African users.
The Lingala and Shona results are a first test, not a finished product. What matters next is whether the winning methods and the data carry over to the rest of the collection and into services people use.




Leave a Reply