Building an ASR Solution for Training and Assessing Children's Reading
This study addresses the scarcity of automatic speech recognition (ASR) systems for African languages—such as Bambara—tailored to children’s read-aloud speech, which hinders reproducible literacy assessments. We present the first open-source ASR benchmark for Bambara child read-aloud data, encompassing field data collection, model adaptation, and classroom validation. Our proposed Soloni model, based on Fast-Conformer and adapted to Bambara phonetics, integrates TDT/CTC decoding with SpecAugment data augmentation, and is benchmarked against QuartzNet. Experimental results reveal architecture-dependent benefits from repeated read-aloud utterances; the optimized Soloni model reduces word error rate (WER) from 0.42 to 0.22 and character error rate (CER) from 0.15 to 0.08, substantially outperforming baseline systems. The model has been successfully deployed in ten classrooms.