Advancing African-Accented Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the performance limitations of automatic speech recognition (ASR) for African-accented English—stemming from severe scarcity of labeled training data—this paper proposes a cognitive uncertainty-driven framework integrating multi-round active learning with core-set selection. Methodologically, we introduce the first Bayesian deep learning–based multi-round adaptive sampling strategy, coupled with a novel Uncertainty-weighted Word Error Rate (U-WER) metric to quantify model adaptability to linguistically challenging accents. Our framework supports fine-tuning of state-of-the-art end-to-end models, including Whisper and Wav2Vec 2.0. Evaluated across multiple African-accented English datasets, the approach achieves an average relative WER reduction of 27% while requiring only 45% of the labeled data needed by baseline methods. It significantly improves generalization to low-resource and out-of-distribution accents. The implementation is publicly available.
📝 Abstract
Accents play a pivotal role in shaping human communication, enhancing our ability to convey and comprehend messages with clarity and cultural nuance. While there has been significant progress in Automatic Speech Recognition (ASR), African-accented English ASR has been understudied due to a lack of training datasets, which are often expensive to create and demand colossal human labor. Combining several active learning paradigms and the core-set approach, we propose a new multi-rounds adaptation process that uses epistemic uncertainty to automate the annotation process, significantly reducing the associated costs and human labor. This novel method streamlines data annotation and strategically selects data samples contributing most to model uncertainty, enhancing training efficiency. We define a new U-WER metric to track model adaptation to hard accents. We evaluate our approach across several domains, datasets, and high-performing speech models. Our results show that our approach leads to a 27% WER relative average improvement while requiring on average 45% less data than established baselines. Our approach also improves out-of-distribution generalization for very low-resource accents, demonstrating its viability for building generalizable ASR models in the context of accented African ASR. We open-source the code here: https://github.com/bonaventuredossou/active_learning_african_asr.
Problem

Research questions and friction points this paper is trying to address.

Improving African-accented English ASR with limited data
Reducing annotation costs via epistemic uncertainty-driven selection
Enhancing model generalization for low-resource accents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Epistemic uncertainty-driven automated annotation process
Multi-rounds adaptation with core-set approach
U-WER metric for tracking accent adaptation
🔎 Similar Papers
No similar papers found.