CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of fine-grained micro-dialect data for Arabic automatic speech recognition (ASR) by constructing a community-crowdsourced speech dataset comprising 40 hours of audio across 21 micro-dialects. Methodologically, it proposes a linguistically grounded micro-dialect labeling taxonomy that reveals sub-regional variations obscured by conventional national labels, while also incorporating code-switching and gender annotations. The framework is systematically evaluated and optimized through multilingual ASR benchmarking, domain adaptation, and audio feature classifiers. Experimental results demonstrate that the optimal system reduces the word error rate from 43.47% to 35.21%, achieving an 85.57% accuracy in micro-dialect identification. These findings provide crucial data resources and methodological foundations for adapting ASR systems to low-resource dialectal settings.
📝 Abstract
We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech spanning 21 micro-dialects across Egypt, Jordan, Lebanon, Mauritania, Palestine, and Saudi Arabia. Each segment is annotated with one or more operational micro-dialect labels, code-switching information, and utterance-level perceived gender, enabling analysis of sub-country variation that is obscured by conventional country-level labels. We describe the collection and annotation process, motivate the micro-dialect inventory linguistically, and benchmark four multilingual ASR systems in zero-shot and adapted settings. The strongest zero-shot system obtains 43.47% aggregate WER, with particularly high error rates on Mauritanian and Lebanese varieties; adaptation on Cardamom reduces its WER to 35.21%. Audio-based identification experiments further show that the annotations provide a learnable prediction target, with a dedicated classifier reaching 85.57% accuracy on 21-way micro-dialect identification. Cardamom provides a resource for studying localized dialectal variation and developing Arabic speech systems with broader regional coverage.
Problem

Research questions and friction points this paper is trying to address.

Arabic ASR
micro-dialectal variation
speech dataset
dialect identification
fine-grained evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Micro-Dialectal Arabic
Speech Dataset
Automatic Speech Recognition
Code-Switching
Dialect Identification
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Bashar Talafha
Bashar Talafha
University of British Columbia
Artificial IntelligenceMachine LearningDeep LearningNatural Language ProcessingAlgorithms
Samar M. Magdy
Samar M. Magdy
The University of British Columbia
LinguisticsComputational LinguisticsNLP
Aisha Alansari
Aisha Alansari
Graduate Assistant, Information and Computer Science Department, KFUPM
Machine LearningNatural Language ProcessingDeep LearningLLMs
A
Alaa Alkhawaldeh
Al al-Bayt University
A
Abdurrahman Juma
Birzeit University
S
Sharaf Makahleh
Jordan University of Science and Technology
N
Nour Gamal
Badr University in Cairo
O
Omar Attia
Badr University in Cairo
H
Hanaa Kurdi
Taibah University
N
Najwa Rizk
Badr University in Cairo
M
Maysa Anaya
H
Hessah Altimyat
Northern Border University
L
Layal Alhazmi
Taibah University
S
Shumukh Alotaibi
Taibah University
H
Hajar Alhadaris
Jordan University of Science and Technology
R
Rayan Alomari
Jordan University of Science and Technology
R
Rahaf Almalaq
Taibah University
M
Malak Alkhorasani
Imam Abdulrahman Bin Faisal University
S
Sara alghamdi
Taibah University
Rahaf Alshamrani
Rahaf Alshamrani
Imam Abdulrahman Bin Faisal University
N
Nsrin Ashraf
El Sewedy University of Technology
I
Ibrahim Jaradat
Jordan University of Science and Technology
N
Nada Qardahji
Jordan University of Science and Technology
Y
Yasmin Zaraket
Imperial College London
E
Elmoukhtar Brahim
Institut Supérieur du Numérique