Projects & Resources

Things I build

My projects connect field-collected language data with reusable corpora, lexical resources, grammatical analysis, and computational models.

Uneme Documentation Project

ELDP-funded · Principal Investigator

A multi-year documentation project creating naturalistic audiovisual recordings, elicitation, annotations, grammatical analyses, lexical resources, and archival deposits focused on Uneme language, cultural practices, oral history, and indigenous knowledge.

View project →

Language documentation ELAR Audiovisual corpus Fieldwork

Speech Corpus & ASR

The research examines the effect of speech-styles, training composition, cross-lingual transfer, and data-efficient adaptation in low-resource ASR.

WhisperXLS-RASRNaturalistic speech

Uneme–English Lexical Resources

A multimedia lexical resource integrating field recordings, elicited lexical data, grammatical information, and corpus examples. The workflow uses FLEx alongside documentary annotations and is designed to support both community-oriented dictionary outputs and linguistic research.

FLExLexicographyCorpus examplesMultimedia

Linguistically Informed ASR Evaluation

Methods for analyzing speech-recognition behavior using language-specific phonological structure, tone, and contrastive features rather than relying only on aggregate word-level scores.

AfricaNLP 2026 paper →

FERToneError analysisAfrican NLP

Grammar & Variation in Edoid Languages

Fieldwork- and corpus-based research on tense, aspect, negation, grammatical tone, focus, movement, and morphosyntactic variation in Uneme and related Edoid languages.

TAMNegationGrammatical toneSyntax

Methods & Tools

Computational

Python PyTorch Hugging Face Whisper XLS-R wav2vec-style models Git/GitHub TensorBoard

Speech & Evaluation

ASR fine-tuning Linguistically-informed analysis Tone-aware evaluation Data preprocessing Cross-lingual transfer

Linguistic Data

ELAN FLEx SayMore Corpus annotation Field recording Lexical databases Archival workflows

Selected Research Outputs

Public datasets, code, archival collections, and publications from my current research program.

Documentation & Archive

Uneme Documentation Project

ELDP-funded documentation of Uneme language, indigenous iron technology, oral history, cultural practices, and linguistic variation across Uneme-speaking communities.

Explore ELAR collection →

Code & Experiments

Uneme ASR Style Generalization

Reproducible experiments investigating naturalistic versus constrained speech, training composition, cross-lingual transfer, and style robustness in low-resource Uneme automatic speech recognition.

View code on GitHub →

Publication

Linguistically Informed Evaluation of Multilingual ASR for African Languages

AfricaNLP 2026 work on Yorùbá and Uneme showing how phonological-feature and tone-aware evaluation can reveal model behavior hidden by aggregate word-level metrics.

Read on ACL Anthology →