AraYoungVoices: A Diverse L1/L2 Corpus of Arabic Child and Adolescent Speech

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inadequate performance of existing automatic speech recognition (ASR) systems for children, adolescents, and second-language (L2) Arabic speakers by constructing the first large-scale multi-dialectal Arabic mixed-speech corpus, comprising 151.7 hours of read speech from 286 native and L2 speakers aged 7–18. Through zero-shot and fine-tuning benchmarks combined with USUP/USSP protocols, mainstream pretrained models are systematically evaluated. The findings reveal differential effects of age and language proficiency on recognition errors: L2 speech poses greater ASR challenges, with the youngest cohort exhibiting the highest error rates. Moreover, joint fine-tuning strategies significantly outperform age-stratified approaches, and ASR outputs demonstrate a tendency to normalize toward standard orthographic forms rather than producing verbatim transcriptions.
📝 Abstract
State-of-the-art ASR systems primarily target native adult speech, leading to substantial performance gaps for children, adolescents, and L2 speakers. We introduce AraYoungVoices, a 151.72-hour Arabic read-speech corpus from 286 speakers aged 7--18, comprising AraKids (7--12) and AraTeens (13--18). The corpus includes 146 native Arabic (L1) and 140 second-language (L2) speakers, with native speakers spanning Egyptian, Gulf, Levantine, and North African dialectal backgrounds and L2 speakers representing diverse linguistic backgrounds across the Americas, Asia, Africa, and Europe. We benchmark four pretrained ASR models under zero-shot and fine-tuned settings using unseen-speaker-$\&$-unseen-prompt (USUP) and unseen-speaker-$\&$-seen-prompt (USSP) evaluations. Results show that L2 speech remains substantially more challenging than L1 speech, with the largest errors observed mainly for younger L2 speakers. Age-specific fine-tuning improves the matched age group, while joint fine-tuning provides a stronger balance across populations. ASR hypotheses are also consistently closer to the standard reading prompt than to the verbatim transcription, particularly for L2 speech, suggesting partial normalization of reading deviations.
Problem

Research questions and friction points this paper is trying to address.

Automatic Speech Recognition
Children and Adolescent Speech
Second Language (L2) Speakers
Arabic Speech Corpus
Innovation

Methods, ideas, or system contributions that make the work stand out.

Arabic Child Speech Corpus
L1/L2 ASR Benchmarking
Age-Specific Fine-Tuning
Dialectal Diversity
Speech Normalization
🔎 Similar Papers
No similar papers found.