Institution profile

RobotsMali

Industry researchafrica · ml
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

Building an ASR Solution for Training and Assessing Children's Reading

Jun 30, 2026

This study addresses the scarcity of automatic speech recognition (ASR) systems for African languages—such as Bambara—tailored to children’s read-aloud speech, which hinders reproducible literacy assessments. We present the first open-source ASR benchmark for Bambara child read-aloud data, encompassing field data collection, model adaptation, and classroom validation. Our proposed Soloni model, based on Fast-Conformer and adapted to Bambara phonetics, integrates TDT/CTC decoding with SpecAugment data augmentation, and is benchmarked against QuartzNet. Experimental results reveal architecture-dependent benefits from repeated read-aloud utterances; the optimized Soloni model reduces word error rate (WER) from 0.42 to 0.22 and character error rate (CER) from 0.15 to 0.08, substantially outperforming baseline systems. The model has been successfully deployed in ten classrooms.

0 citationsRead paper

Generative Artificial Intelligence, Musical Heritage and the Construction of Peace Narratives: A Case Study in Mali

Jan 21, 2026

This work proposes a participatory framework embedding generative artificial intelligence within cultural reflexivity to address ethnic tensions and cultural fragmentation in Mali. By fostering human-AI co-creation that integrates local linguistic corpora and traditional musical elements, the approach supports the reconstruction of indigenous musical narratives while upholding cultural authenticity. The methodology amplifies local voices, strengthens cultural sovereignty, and fosters social cohesion through a pathway of “symbolic diplomacy” rather than cultural homogenization. Empirical implementation reveals critical challenges—including data scarcity, algorithmic transparency, and copyright ethics—offering both an innovative paradigm and reflective insights for AI-driven revitalization of intangible cultural heritage.

0 citationsRead paper

Listen, Attend, Understand: a Regularization Technique for Stable E2E Speech Translation Training on High Variance labels

Jan 03, 2026arXiv.org

This work addresses the instability and performance limitations of end-to-end speech translation when target transcriptions exhibit high variance and semantic ambiguity. The authors propose LAU, a method that introduces a directional auxiliary loss by freezing textual embeddings to impose semantic regularization on the acoustic encoder’s latent space, prioritizing semantic fidelity over literal phonetic details. To further enforce this constraint, the encoder’s weight structure is reconfigured. Additionally, the Total Parameter Drift metric is introduced to quantify the impact of regularization. Evaluated on only 30 hours of non-expertly annotated Bambara–French data, the LAU model achieves or surpasses the performance of systems pretrained with twice the amount of data, demonstrating significant improvements in both semantic fidelity and training stability.

0 citationsRead paper

Kunnafonidilaw ka Cadeau: an ASR dataset of present-day Bambara

Dec 22, 2025

Modern Bambara lacks authentic, real-world automatic speech recognition (ASR) datasets, hindering robust model development. Method: We introduce Kunkado, a 160-hour spontaneous speech dataset derived from Malian broadcast audio, the first to systematically encompass code-switching, disfluencies, background noise, and speaker overlap. We propose a transcription normalization framework addressing non-standard expressions—including numeric forms, multilingual tags, and pause markers—and fine-tune the Parakeet ASR model using a human-verified subset. Contribution/Results: Our approach reduces word error rate (WER) by 7.35% and 3.74% on two real-world test sets, respectively, significantly outperforming a baseline model trained solely on 98 hours of clean speech with identical architecture, as confirmed by human evaluation. All data, annotations, and models are publicly released, establishing a new benchmark and practical paradigm for robust ASR in low-resource languages.

0 citationsRead paper

Dealing with the Hard Facts of Low-Resource African NLP

Nov 23, 2025

Low-resource African languages—such as Bambara—face critical bottlenecks in speech technology development due to severe data scarcity, high annotation costs, and the absence of standardized evaluation frameworks. To address these challenges, this work proposes an end-to-end solution: (1) field-collecting 612 hours of spontaneous, conversational Bambara speech; (2) designing a semi-automated annotation pipeline with multi-tier human verification; (3) leveraging self-supervised speech representation learning to build ultra-compact monolingual models; and (4) establishing a trustworthy evaluation framework integrating automated metrics with expert-led human assessment. Key contributions include: the first large-scale, open-source Bambara speech dataset; a series of lightweight pre-trained models optimized for low-resource settings; and a comprehensive evaluation toolkit. Empirical results demonstrate substantial improvements in feasibility, reproducibility, and real-world deployability of automatic speech recognition for under-resourced African languages.

0 citationsRead paper
Recent publications

Latest Papers

Building an ASR Solution for Training and Assessing Children's Reading

Jun 30, 2026

This study addresses the scarcity of automatic speech recognition (ASR) systems for African languages—such as Bambara—tailored to children’s read-aloud speech, which hinders reproducible literacy assessments. We present the first open-source ASR benchmark for Bambara child read-aloud data, encompassing field data collection, model adaptation, and classroom validation. Our proposed Soloni model, based on Fast-Conformer and adapted to Bambara phonetics, integrates TDT/CTC decoding with SpecAugment data augmentation, and is benchmarked against QuartzNet. Experimental results reveal architecture-dependent benefits from repeated read-aloud utterances; the optimized Soloni model reduces word error rate (WER) from 0.42 to 0.22 and character error rate (CER) from 0.15 to 0.08, substantially outperforming baseline systems. The model has been successfully deployed in ten classrooms.

0 citationsRead paper

Generative Artificial Intelligence, Musical Heritage and the Construction of Peace Narratives: A Case Study in Mali

Jan 21, 2026

This work proposes a participatory framework embedding generative artificial intelligence within cultural reflexivity to address ethnic tensions and cultural fragmentation in Mali. By fostering human-AI co-creation that integrates local linguistic corpora and traditional musical elements, the approach supports the reconstruction of indigenous musical narratives while upholding cultural authenticity. The methodology amplifies local voices, strengthens cultural sovereignty, and fosters social cohesion through a pathway of “symbolic diplomacy” rather than cultural homogenization. Empirical implementation reveals critical challenges—including data scarcity, algorithmic transparency, and copyright ethics—offering both an innovative paradigm and reflective insights for AI-driven revitalization of intangible cultural heritage.

0 citationsRead paper

Listen, Attend, Understand: a Regularization Technique for Stable E2E Speech Translation Training on High Variance labels

Jan 03, 2026arXiv.org

This work addresses the instability and performance limitations of end-to-end speech translation when target transcriptions exhibit high variance and semantic ambiguity. The authors propose LAU, a method that introduces a directional auxiliary loss by freezing textual embeddings to impose semantic regularization on the acoustic encoder’s latent space, prioritizing semantic fidelity over literal phonetic details. To further enforce this constraint, the encoder’s weight structure is reconfigured. Additionally, the Total Parameter Drift metric is introduced to quantify the impact of regularization. Evaluated on only 30 hours of non-expertly annotated Bambara–French data, the LAU model achieves or surpasses the performance of systems pretrained with twice the amount of data, demonstrating significant improvements in both semantic fidelity and training stability.

0 citationsRead paper

Kunnafonidilaw ka Cadeau: an ASR dataset of present-day Bambara

Dec 22, 2025

Modern Bambara lacks authentic, real-world automatic speech recognition (ASR) datasets, hindering robust model development. Method: We introduce Kunkado, a 160-hour spontaneous speech dataset derived from Malian broadcast audio, the first to systematically encompass code-switching, disfluencies, background noise, and speaker overlap. We propose a transcription normalization framework addressing non-standard expressions—including numeric forms, multilingual tags, and pause markers—and fine-tune the Parakeet ASR model using a human-verified subset. Contribution/Results: Our approach reduces word error rate (WER) by 7.35% and 3.74% on two real-world test sets, respectively, significantly outperforming a baseline model trained solely on 98 hours of clean speech with identical architecture, as confirmed by human evaluation. All data, annotations, and models are publicly released, establishing a new benchmark and practical paradigm for robust ASR in low-resource languages.

0 citationsRead paper

Dealing with the Hard Facts of Low-Resource African NLP

Nov 23, 2025

Low-resource African languages—such as Bambara—face critical bottlenecks in speech technology development due to severe data scarcity, high annotation costs, and the absence of standardized evaluation frameworks. To address these challenges, this work proposes an end-to-end solution: (1) field-collecting 612 hours of spontaneous, conversational Bambara speech; (2) designing a semi-automated annotation pipeline with multi-tier human verification; (3) leveraging self-supervised speech representation learning to build ultra-compact monolingual models; and (4) establishing a trustworthy evaluation framework integrating automated metrics with expert-led human assessment. Key contributions include: the first large-scale, open-source Bambara speech dataset; a series of lightweight pre-trained models optimized for low-resource settings; and a comprehensive evaluation toolkit. Empirical results demonstrate substantial improvements in feasibility, reproducibility, and real-world deployability of automatic speech recognition for under-resourced African languages.

0 citationsRead paper