
Research and open source
Open, verifiable science
We publish our methods, our datasets, and our models so the scientific community and developers can check them, challenge them, and build on them.
01
Scientific publications
02
Open source projects
The code and corpora that make our models possible are published on GitHub, most of them under the MIT license.
NMTMD
Open Ewe-language corpus for neural machine translation in West Africa.
MIT · 26 stars · 10 forksumini_speech
Official training repository for Yodi V1, the first speech recognition system for Ewe.
Research · speech recognitionYodi
Speech recognition model for 8 words in Ewe, a Python package for fast inference.
MIT · PythonTogo-Data-Lab-Masterclass
Teaching notebooks: Transformer architecture and its mathematical foundations.
Jupyter · community03
Models and datasets
Umbaji publishes its speech synthesis models and corpora directly on Hugging Face, for researchers and developers to reuse.