Research

Language structure, primary data, and low-resource AI

My research asks how carefully documented linguistic data can improve both our understanding of language and the technologies built for languages that remain underrepresented in modern AI.

Research program

Four connected research directions

My work connects documentary fieldwork, linguistic analysis, corpus development, and computational modeling. These areas are not separate projects; they form a common research program focused on how underdescribed languages can inform both linguistic theory and language technology.

01 · Speech Technology

Low-Resource Speech Technology

How can speech systems learn robustly from small, heterogeneous, and linguistically complex datasets?

I build and evaluate automatic speech recognition systems for languages with limited transcribed speech resources. My current work investigates cross-lingual transfer, data-efficient adaptation, speech-style mismatch, naturalistic versus constrained speech, tonal contrasts, and the limits of multilingual pretraining.

A central concern is ecological validity: models that perform well on carefully elicited or read speech may behave very differently on spontaneous documentary recordings. I therefore treat recording style, speaker variation, community variation, and linguistic structure as central experimental variables rather than noise to be removed.

ASR Whisper XLS-R Cross-lingual transfer Naturalistic speech Style robustness

02 · Model Evaluation

Linguistically Informed AI Evaluation

What do aggregate metrics hide about what speech and language models actually learn?

Word Error Rate and Character Error Rate are useful, but they can collapse linguistically different error types into a single score. Our work develops and applies evaluation approaches that ask which phonological, tonal, morphological, and grammatical distinctions are preserved or lost by a model.

This direction is represented by our recent AfricaNLP 2026 work on multilingual ASR for African languages, where phonological-feature and tone-sensitive evaluation reveals systematic model behavior that word-level accuracy alone cannot show.

Read the AfricaNLP 2026 paper →

WER / CER Phonological features Tone Error analysis African NLP

03 · Grammar & Variation

Corpus-Accountable Grammar & Variation

How can grammatical description capture the variation that actually occurs in documentary corpora?

I study tense, aspect, negation, grammatical tone, focus, movement, and clause structure in underdescribed African languages, especially Edoid languages.

I am particularly interested in the relationship between targeted elicitation and naturalistic evidence, and in grammatical descriptions that remain accountable to patterns across speakers, communities, varieties, and discourse contexts.

My work treats variation not as a defect in the dataset, but as an empirical fact with consequences for linguistic theory, documentation practice, and computational modeling.

Tense & aspect Negation Grammatical tone Focus Clause structure Variation

04 · Research Infrastructure

Documentation-Informed Language Technology

How can documentary resources become computationally useful without stripping away linguistic richness?

For many underdescribed languages, computational research cannot begin with a large downloaded benchmark. It begins with field relationships, recording decisions, transcription conventions, annotation, lexical analysis, and language-specific knowledge.

I therefore work across the full pipeline: audiovisual documentation, transcription and annotation, corpus and lexicon development, grammatical analysis, speech-dataset creation, model training, and linguistic evaluation.

This creates a feedback loop in which documentation informs technology, while model errors generate new questions about language structure, representation, and variation.

Fieldwork Corpus development ELAN FLEx Speech datasets Model evaluation

Research principle

Documentation and technology should inform one another

Language documentation should not end with an archive, and language technology should not begin with a downloaded dataset.

For underdescribed languages, high-quality technology depends on understanding how the data was collected, how the language is structured, how speakers and varieties differ, and which linguistic distinctions matter.

Empirical focus

Uneme as a research laboratory

Much of my current work centers on Uneme, an Edoid language of Nigeria, alongside comparative and grammatical research on related Edoid languages and computational evaluation involving Yorùbá.

Uneme provides an unusually rich setting for connecting fieldwork, documentation, grammatical analysis, variation, speech technology, and model evaluation within a single research program.

The goal, however, is broader than a single language: I use Uneme to develop methods and questions that can generalize to other underdescribed and low-resource languages.

Primary empirical language

Uneme

Comparative focus

Edoid languages · Yorùbá

Research domains

Documentation · Grammar · Variation · ASR · Linguistically informed evaluation