Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German

📅 2026-05-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient evaluation of commercial speech recognition systems in code-switching scenarios involving Arabic, Persian, and German with English, where conventional word error rate (WER) fails to capture semantic accuracy amid transcriptional discrepancies. To bridge this gap, the authors construct a high-quality multilingual code-switching benchmark and propose an efficient data curation pipeline that combines heuristic rules with large language models (GPT-4o and Gemini 1.5 Pro), substantially reducing annotation costs. They introduce BERTScore as a semantic evaluation metric and employ embedding projection and structural signal analysis to uncover performance stratification across difficulty levels. Evaluations reveal that ElevenLabs Scribe v2 achieves the best performance (WER: 13.2%, BERTScore: 0.936). The work also releases a dataset of 1,200 annotated samples, demonstrating that semantic similarity can be preserved independently of surface-level textual variations.
📝 Abstract
Code-switching -- the natural alternation between two languages within a single utterance -- represents one of the most challenging and under-studied conditions for automatic speech recognition (ASR). Existing commercial ASR benchmarks predominantly evaluate clean, monolingual audio and report a single Word Error Rate (WER) figure that tells practitioners little about real-world multilingual performance. We present a benchmark evaluating five commercial ASR providers across four language pairs: Egyptian Arabic--English, Saudi Arabic (Najdi/Hijazi)--English, Persian (Farsi)--English, and German--English. Each dataset comprises 300 samples selected by a two-stage pipeline: a heuristic filter scoring transcripts on five structural code-switching signals, followed by a GPT-4o and Gemini 1.5 Pro ensemble scoring candidates across six linguistic dimensions. This pipeline reduces LLM scoring costs by approximately 91\% relative to exhaustive scoring. We evaluate the systems on both WER and BERTScore, arguing that BERTScore is a more reliable metric for Arabic and Persian pairs where transliteration variance causes WER to penalise semantically correct transcriptions. ElevenLabs Scribe v2 achieves the lowest WER across all four language pairs (13.2% overall; 13.1% on Egyptian Arabic) and leads on BERTScore (0.936 overall). We further demonstrate that difficulty-stratified analysis reveals performance gaps masked by aggregate averages, and that BERT embedding projections confirm semantic proximity between reference and hypothesis despite surface-level script differences. The benchmarking dataset is publicly available at https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch.
Problem

Research questions and friction points this paper is trying to address.

code-switching
automatic speech recognition
benchmarking
multilingual speech
Word Error Rate
Innovation

Methods, ideas, or system contributions that make the work stand out.

code-switching
ASR benchmarking
BERTScore
multilingual speech recognition
cost-efficient LLM filtering
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Sajjad Abdoli
Sajjad Abdoli
perle.ai
Deep LearningMusic Information RetrievalAdversarial Machine Learning
G
Ghassan Al-Sumaidaee
Perle AI
C
Clayton W. Taylor
Perle AI
A
Ahmad ElShiekh
Perle AI
A
Ahmed Rashad
Perle AI