🤖 AI Summary
Modern Bambara lacks authentic, real-world automatic speech recognition (ASR) datasets, hindering robust model development. Method: We introduce Kunkado, a 160-hour spontaneous speech dataset derived from Malian broadcast audio, the first to systematically encompass code-switching, disfluencies, background noise, and speaker overlap. We propose a transcription normalization framework addressing non-standard expressions—including numeric forms, multilingual tags, and pause markers—and fine-tune the Parakeet ASR model using a human-verified subset. Contribution/Results: Our approach reduces word error rate (WER) by 7.35% and 3.74% on two real-world test sets, respectively, significantly outperforming a baseline model trained solely on 98 hours of clean speech with identical architecture, as confirmed by human evaluation. All data, annotations, and models are publicly released, establishing a new benchmark and practical paradigm for robust ASR in low-resource languages.
📝 Abstract
We present Kunkado, a 160-hour Bambara ASR dataset compiled from Malian radio archives to capture present-day spontaneous speech across a wide range of topics. It includes code-switching, disfluencies, background noise, and overlapping speakers that practical ASR systems encounter in real-world use. We finetuned Parakeet-based models on a 33.47-hour human-reviewed subset and apply pragmatic transcript normalization to reduce variability in number formatting, tags, and code-switching annotations. Evaluated on two real-world test sets, finetuning with Kunkado reduces WER from 44.47% to 37.12% on one and from 36.07% to 32.33% on the other. In human evaluation, the resulting model also outperforms a comparable system with the same architecture trained on 98 hours of cleaner, less realistic speech. We release the data and models to support robust ASR for predominantly oral languages.