Kunnafonidilaw ka Cadeau: an ASR dataset of present-day Bambara

📅 2025-12-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Modern Bambara lacks authentic, real-world automatic speech recognition (ASR) datasets, hindering robust model development. Method: We introduce Kunkado, a 160-hour spontaneous speech dataset derived from Malian broadcast audio, the first to systematically encompass code-switching, disfluencies, background noise, and speaker overlap. We propose a transcription normalization framework addressing non-standard expressions—including numeric forms, multilingual tags, and pause markers—and fine-tune the Parakeet ASR model using a human-verified subset. Contribution/Results: Our approach reduces word error rate (WER) by 7.35% and 3.74% on two real-world test sets, respectively, significantly outperforming a baseline model trained solely on 98 hours of clean speech with identical architecture, as confirmed by human evaluation. All data, annotations, and models are publicly released, establishing a new benchmark and practical paradigm for robust ASR in low-resource languages.

Technology Category

Natural Language Processing: SpeechMachine Learning: Large Multimodal Models (LMMs)Data Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
We present Kunkado, a 160-hour Bambara ASR dataset compiled from Malian radio archives to capture present-day spontaneous speech across a wide range of topics. It includes code-switching, disfluencies, background noise, and overlapping speakers that practical ASR systems encounter in real-world use. We finetuned Parakeet-based models on a 33.47-hour human-reviewed subset and apply pragmatic transcript normalization to reduce variability in number formatting, tags, and code-switching annotations. Evaluated on two real-world test sets, finetuning with Kunkado reduces WER from 44.47% to 37.12% on one and from 36.07% to 32.33% on the other. In human evaluation, the resulting model also outperforms a comparable system with the same architecture trained on 98 hours of cleaner, less realistic speech. We release the data and models to support robust ASR for predominantly oral languages.
Problem

Research questions and friction points this paper is trying to address.

Develops a Bambara ASR dataset from radio archives
Addresses real-world speech challenges like code-switching and noise
Improves ASR accuracy for oral languages via finetuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compiled 160-hour Bambara ASR dataset from radio archives
Finetuned Parakeet-based models with pragmatic transcript normalization
Achieved lower WER on real-world test sets through finetuning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yacouba Diarra
RobotsMali AI4D Lab, Bamako, Mali
P
Panga Azazia Kamate
RobotsMali AI4D Lab, Bamako, Mali
N
Nouhoum Souleymane Coulibaly
RobotsMali AI4D Lab, Bamako, Mali
Michael Leventhal
Michael Leventhal
RobotsMali
machine learninglow-resource languagesparallel computing architecturesautomata processingXML