Giving a voice to languages without tools
In 2023, Ewè had no public speech recognition system. Yodi was built to fill that gap: recognizing Ewè speech is the first step toward automatic transcription, translation and voice assistants in Togolese languages.
System architecture
The system pairs word-to-waveform mapping with convolutional neural network classifiers. The acoustic signal is segmented into units matching target words, and each segment is classified, avoiding dependence on the large pre-trained models that do not exist for the language.
A data problem above all
The main obstacle was the lack of annotated audio data. Recordings were collected from native speakers, transcripts were verified by annotators, and the result was structured into a reference corpus for training and evaluation.
Limits and legacy
The first version covers a restricted vocabulary and a controlled recording context. It nonetheless establishes the methodology — collection, annotation, evaluation — that later Yodi releases take up and generalize to other languages and open domains.