01
Data collection
Voice and text data are gathered from native speakers, public sources, and curated multilingual content to build a broad and representative foundation.
Research · Publications
Our research write-ups on automatic speech recognition, neural text-to-speech, and machine translation for low-resource African languages — along with the data pipeline that makes them possible.
Data pipeline
Every model begins with a pipeline that is designed to be repeatable, transparent, and practical for low-resource languages. The process spans collection, cleaning, annotation, evaluation, and release.
01
Voice and text data are gathered from native speakers, public sources, and curated multilingual content to build a broad and representative foundation.
02
Audio is normalized, segmented, and quality-checked while text is standardized and aligned to reduce noise and improve learning quality.
03
Transcripts, translations, and metadata are verified so the data is reliable enough for training and evaluation.
04
Models are fine-tuned, validated, and benchmarked with both automated metrics and human review to measure real-world usefulness.
05
The best-performing models and datasets are packaged for reuse, while feedback from deployment helps shape the next iteration.
Publications
These publications capture the research direction behind the data pipeline, model development, and language infrastructure work.
A detailed study of the development of a Krio speech-to-text system, including data collection, preprocessing, model training, and evaluation.
Read more →An exploration of neural TTS methods for Krio, focused on naturalness, data preparation, and extension to other African languages.
Read more →A multilingual translation study beginning with Krio and outlining a broader path for African language translation systems.
Read more →