All publications

research

YodiV3: NLP for Togolese Languages with the Eyaa-Tom Dataset and the Lom Metric

Umbaji research team

Presents Eyaa-Tom, a multi-domain parallel corpus for ten Togolese languages and two lingua francas, and introduces the Lom metric to quantify each language's AI readiness.

Published on 2/18/2026 · AfricaNLP Workshop, ACL Anthology · 1 min read

Why Togolese languages are absent from AI

Current language models ignore almost every West African language because structured corpora do not exist. YodiV3 tackles this at the root: by documenting ten Togolese languages in a public corpus, then by measuring mathematically where each one stands on the road to AI.

Eyaa-Tom, a corpus built to be real

Eyaa-Tom is a multi-domain parallel corpus: the text comes not from machine-translated web pages but from conversations, administrations and healthcare situations actually collected in the field. A community of native speakers produces and validates every sentence pair.

  • Ten Togolese languages covered, plus two reference lingua francas
  • Everyday usage domains, not only administrative writing
  • Double human validation before publication

The Lom metric

Lom quantifies a language's AI readiness: available data volume, domain coverage, availability of parallel resources, and the capacity to train useful models. Rather than a fixed ranking, Lom gives each language an actionable diagnosis to decide what to collect next.

Results and next steps

Early models trained on Eyaa-Tom show that transcription and translation remain possible even with modest volumes, as long as quality and domain are controlled. The corpus and the Lom scores are published openly, and the next iterations target the least-resourced languages.

CitationUmbaji research team 2026. “YodiV3: NLP for Togolese Languages with the Eyaa-Tom Dataset and the Lom Metric”. AfricaNLP Workshop, ACL Anthology
Read the publication