Benchmarking Automatic Speech Recognition Tools for Iberian Languages

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of automatic speech recognition (ASR) evaluations for Iberian languages and the unclear trade-off between bias and efficiency in low-resource settings. To this end, we construct an 85-hour multi-scenario dataset and systematically benchmark eleven ASR systems across five Iberian languages. Employing word error rate (WER) and real-time factor (RTF) as evaluation metrics, we comprehensively compare the performance of open-source models against commercial APIs. Our results demonstrate that no single optimal model exists, revealing a complex trade-off between recognition accuracy and inference efficiency. This work fills a critical gap in ASR evaluation for these languages and provides empirical guidance for model selection in practical deployments.
📝 Abstract
Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Catalan, Galician, Portuguese, Spanish), with German and Turkish as controls. Evaluation uses an 85-hour dataset covering read speech, broadcast media, and audiobooks, assessing accuracy and efficiency via word error rate (WER) and real-time factors (RTF/RTFx). Results show no single model dominates: accuracy, efficiency, and language coverage present clear trade-offs. Low-resource languages, especially Basque, degrade significantly, highlighting the role of training coverage. We observe consistent sex disparities across most systems, highlighting fairness challenges in multilingual ASR. Overall, the benchmark provides practical guidance for real-world model selection.
Problem

Research questions and friction points this paper is trying to address.

Automatic Speech Recognition
Iberian Languages
Low-resource Languages
Benchmarking
Fairness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automatic Speech Recognition
Benchmarking
Iberian Languages
Low-resource Languages
Fairness
🔎 Similar Papers
No similar papers found.