Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the neglect of dialectal variation and the absence of culturally grounded benchmarks in Arabic speech technology by constructing a culturally situated benchmark encompassing 17 dialects, including the first Ahwazi Arabic test set. Multimodal data were collected by having native speakers describe images, and mainstream automatic speech recognition (ASR) systems, such as Whisper and GPT-4o, were evaluated in a zero-shot setting using jointly reported Word Error Rate (WER) and Character Error Rate (CER) metrics. Results reveal substantial performance disparities across dialects, with the best WER still reaching 35%, exposing critical bottlenecks in current dialect recognition capabilities. This work establishes an essential evaluation foundation for low-resource Arabic speech research.
📝 Abstract
Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unseen by existing models. Dialect-community coordinators selected culturally relevant images across 10 topics, and native speakers described them through five structured scenarios, yielding approximately 50 minutes per dialect (13.7 hours total). We benchmark 12 state-of-the-art ASR systems zero-shot, including GPT-4o-transcribe, Voxtral-Mini-4B, Fanar-STT-LF, Whisper, SeamlessM4T-v2, and wav2vec2-based models. GPT-4o-transcribe achieves the lowest overall WER at 35.0%, followed by Voxtral-Mini-4B, Fanar-STT-LF, and Whisper-Large-v3 at 41.1%, 45.9%, and 49.5%, respectively, indicating substantial remaining errors across Arabic dialect communities. Performance varies considerably across dialect groups, with no model performing uniformly best across all groups. WER alone also obscures dialectal ASR behaviour: wav2vec2-based models show large WER/CER gaps, where character-level agreement remains much higher than word-level accuracy, motivating joint WER/CER reporting. ALMIEYAR provides a unified benchmark for culturally grounded Arabic ASR evaluation, including the first published benchmark for Ahwazi Arabic.
Problem

Research questions and friction points this paper is trying to address.

Arabic dialects
speech recognition
ASR benchmark
cultural grounding
dialectal evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-dialect Arabic ASR
Culturally grounded benchmark
Zero-shot evaluation
WER/CER analysis
Low-resource speech recognition
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Omid Ghahroodi
Omid Ghahroodi
Research Assistant at Qatar Computing Research Institute, Sharif University of Technology Alumni
Machine LearningDeep LearningNatural Language ProcessingLLMVLM
A
Anas Madkoor
QCRI, HBKU
D
Dima Faris Al Saudi
QCRI, HBKU
F
Fagr Tahir
Qatar University
M
Malak Annan
QCRI, HBKU
T
Talha shahid javad allah rakha
UDST
O
Omar Al-Busaidi
Algo AI
Z
Zineb El Kahla
Qatar University
I
Iheb Zouari
QCRI, HBKU
E
Essa Ahmed Abou Jabal
Qatar University
A
Ahmed Ezzat
Qatar University
H
Hind AL-Merekhi
QCRI, HBKU
A
Aisha Hamad M A Al-Naimi
Qatar University
H
Hadi Wazni
UCL
B
Bushra Alnajjar
Qatar University
O
Omar Amin
Qatar University
H
Haya Al-Thani
QCRI, HBKU
H
Houssam Eddine-Othman Lachemat
QCRI, HBKU
M
Marwa Elwakedy
University of Tripoli
S
Sundus Abdulmalik Al Nahari
AUC
E
Elahe Zahiri
QCRI, HBKU
O
Osamah Sarraj
KAUST
R
Raghad Mousa
CMU-Q
M
Mckeen Assi
QCRI, HBKU
A
Ahd Al Jumah
KFUPM