Low-Resource Speech Technology
ASR, cross-lingual transfer, data-efficient adaptation, naturalistic speech, style mismatch, and robust evaluation for languages with limited training data.
Linguistics · Speech Technology · African Languages
PhD Researcher in Linguistics & Low-Resource Speech & Language Technology
I build datasets, linguistic analyses, and computational tools for underdescribed African languages. My work connects primary fieldwork with corpus development, grammatical analysis, low-resource speech modeling, and linguistically informed AI evaluation.
Research program
I work across a research pipeline that is often split across different disciplines: collecting and documenting language data, turning it into structured corpora, analyzing linguistic systems and variation, training computational models, and using linguistic knowledge to understand model behavior.
ASR, cross-lingual transfer, data-efficient adaptation, naturalistic speech, style mismatch, and robust evaluation for languages with limited training data.
Community-based fieldwork, audiovisual documentation, annotation, corpus building, lexical resources, multilingual dictionaries, and archival workflows designed for both linguistic and computational reuse.
Tense, aspect, negation, grammatical tone, clause structure, variation, and model-error analysis grounded in language-specific phonology and grammar.
Featured publication
Standard aggregate metrics can obscure what speech systems actually learn about low-resource tonal languages. This work evaluates multilingual ASR on Yorùbá and Uneme with character-, phonological-feature-, and tone-sensitive measures, revealing systematic model behavior hidden by word-level accuracy alone.
Featured project
I lead an ELDP-supported, multi-year documentation project on Uneme, an Edoid language of Nigeria. The project combines naturalistic audiovisual documentation with corpus development, grammatical analysis, lexical resources, archival outputs, and speech technology.
Naturalistic discourse, oral histories, cultural practices, traditional knowledge, material culture, and targeted linguistic elicitation.
Transcription, annotation, grammatical analysis, variation across communities, and an Uneme–English multimedia lexical resource.
Speech datasets and low-resource ASR experiments that test how modern models handle naturalistic, tonal, and highly constrained-data conditions.
Recent highlights
Scholarly Engagement
Recent presentations spanning low-resource speech technology, African language structure, grammatical theory, and language documentation.
Broader vision
African languages shouldn't remain passive "low-resource benchmarks". They should actively contribute to both linguistic theory and the design of next-generation language technology.
As a long-term goal, I want to develop research approaches in which language documentation, linguistic theory, corpus development, and machine learning reinforce one another. I am especially interested in technologies that work with the realities of underdescribed languages— limited data, naturalistic speech, tonal and morphological complexity, and speaker and dialect variation.
Collaboration
I welcome conversations with researchers, language communities, research labs, and organizations. I am especially interested in collaborations that connect data, language structure, and computational modeling in ways that expand both scientific understanding and technological access.