🤖 AI Summary
This study addresses the scarcity of fine-grained micro-dialect data for Arabic automatic speech recognition (ASR) by constructing a community-crowdsourced speech dataset comprising 40 hours of audio across 21 micro-dialects. Methodologically, it proposes a linguistically grounded micro-dialect labeling taxonomy that reveals sub-regional variations obscured by conventional national labels, while also incorporating code-switching and gender annotations. The framework is systematically evaluated and optimized through multilingual ASR benchmarking, domain adaptation, and audio feature classifiers. Experimental results demonstrate that the optimal system reduces the word error rate from 43.47% to 35.21%, achieving an 85.57% accuracy in micro-dialect identification. These findings provide crucial data resources and methodological foundations for adapting ASR systems to low-resource dialectal settings.
📝 Abstract
We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech spanning 21 micro-dialects across Egypt, Jordan, Lebanon, Mauritania, Palestine, and Saudi Arabia. Each segment is annotated with one or more operational micro-dialect labels, code-switching information, and utterance-level perceived gender, enabling analysis of sub-country variation that is obscured by conventional country-level labels. We describe the collection and annotation process, motivate the micro-dialect inventory linguistically, and benchmark four multilingual ASR systems in zero-shot and adapted settings. The strongest zero-shot system obtains 43.47% aggregate WER, with particularly high error rates on Mauritanian and Lebanese varieties; adaptation on Cardamom reduces its WER to 35.21%. Audio-based identification experiments further show that the annotations provide a learnable prediction target, with a dedicated classifier reaching 85.57% accuracy on 21-way micro-dialect identification. Cardamom provides a resource for studying localized dialectal variation and developing Arabic speech systems with broader regional coverage.