SHAMS: An Audio-Grounded Pronunciation Benchmark for Levantine Arabic

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the evaluation challenges in speech technologies for Levantine Arabic, which arise from dialectal diversity and non-standard orthography. To this end, we propose the first audio-anchored, multi-dialectal hierarchical benchmark. Comprising 1,300 audio recordings with four tiers of aligned data, this benchmark integrates multimodal alignment, phonetic transcription, automatic speech recognition (ASR), and grapheme-to-phoneme conversion to overcome orthographic opacity and facilitate the evaluation of diverse downstream tasks. Through systematic experiments, we assess the performance of both open-source and proprietary models. Ultimately, this work establishes a standardized and reliable evaluation paradigm for measuring advancements in speech technologies for Levantine Arabic.
📝 Abstract
Levantine Arabic (LA) is spoken by tens of millions of people, creating a pressing need for shared benchmarks to evaluate LA speech-language technologies. Evaluating such technology is particularly challenging given LA's internal diversity and its opaque and non-standardized orthography. We present SHAMS (SHami Annotated Multi-dialect Speech), a benchmark comprising 1,300 utterances drawn from open audio corpora, balanced across five LA varieties (Urban and Rural Palestinian, and Urban Jordanian, Lebanese, and Syrian). Each utterance is represented across four aligned tiers: audio, unvocalized orthography, diacritized text, and phonetic transcription. This structure supports evaluation of various downstream tasks such as diacritization, grapheme-to-phoneme conversion, automatic speech recognition, and audio-to-phoneme, grounded in audio and stratified by variety. We benchmark open and proprietary models across these tasks to demonstrate the utility of this benchmark for measuring progress across LA. We release SHAMS at https://shams-nlp.github.io .
Problem

Research questions and friction points this paper is trying to address.

Levantine Arabic
pronunciation benchmark
speech-language technology
dialect diversity
orthography
Innovation

Methods, ideas, or system contributions that make the work stand out.

Levantine Arabic
Pronunciation Benchmark
Multi-dialect Speech
Phonetic Transcription
Speech-Language Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Ben Sapirstein
Reichman University
R
Roy Mattar
Independent Researcher
G
Guy Mor-Lan
The Hebrew University of Jerusalem
A
Ahlam Mohamed
Tel Aviv University
L
Letizia Cerqueglini
Tel Aviv University
Morris Alper
Morris Alper
Machine Learning Researcher
machine learningcomputational linguisticsnatural language processingmultimodal learning