🤖 AI Summary
This study addresses the inadequate performance of existing automatic speech recognition (ASR) systems for children, adolescents, and second-language (L2) Arabic speakers by constructing the first large-scale multi-dialectal Arabic mixed-speech corpus, comprising 151.7 hours of read speech from 286 native and L2 speakers aged 7–18. Through zero-shot and fine-tuning benchmarks combined with USUP/USSP protocols, mainstream pretrained models are systematically evaluated. The findings reveal differential effects of age and language proficiency on recognition errors: L2 speech poses greater ASR challenges, with the youngest cohort exhibiting the highest error rates. Moreover, joint fine-tuning strategies significantly outperform age-stratified approaches, and ASR outputs demonstrate a tendency to normalize toward standard orthographic forms rather than producing verbatim transcriptions.
📝 Abstract
State-of-the-art ASR systems primarily target native adult speech, leading to substantial performance gaps for children, adolescents, and L2 speakers. We introduce AraYoungVoices, a 151.72-hour Arabic read-speech corpus from 286 speakers aged 7--18, comprising AraKids (7--12) and AraTeens (13--18). The corpus includes 146 native Arabic (L1) and 140 second-language (L2) speakers, with native speakers spanning Egyptian, Gulf, Levantine, and North African dialectal backgrounds and L2 speakers representing diverse linguistic backgrounds across the Americas, Asia, Africa, and Europe. We benchmark four pretrained ASR models under zero-shot and fine-tuned settings using unseen-speaker-$\&$-unseen-prompt (USUP) and unseen-speaker-$\&$-seen-prompt (USSP) evaluations. Results show that L2 speech remains substantially more challenging than L1 speech, with the largest errors observed mainly for younger L2 speakers. Age-specific fine-tuning improves the matched age group, while joint fine-tuning provides a stronger balance across populations. ASR hypotheses are also consistently closer to the standard reading prompt than to the verbatim transcription, particularly for L2 speech, suggesting partial normalization of reading deviations.