Vimarsha: Faithful ASR Evaluation for Indian Languages with Demographic Diversity, In-the-Wild Audio and Spelling Variations

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入Vimarsha基准来解决印度语言ASR评估中的系统偏差问题,该基准结合了多样化的现场录音、野生音频以及多变体转录框架。
📝 Abstract
Evaluation benchmarks for Indian language automatic speech recognition (ASR) suffer from two systematic biases: optimistic scores from clean, controlled audio conditions, and pessimistic scores from overly rigid transcription standards that penalize valid linguistic variations. We introduce Vimarsha, a 100-hour benchmark spanning all 22 scheduled Indian languages, designed to address both distortions. Vimarsha combines demographically diverse on-field recordings with carefully mined in-the-wild audio selected for acoustic difficulty, alongside a lattice of variations framework that encodes multiple valid transcriptions per utterance. Evaluations of 10 state-of-the-art ASR models reveal substantial shifts in model rankings under realistic conditions, geographic and demographic performance disparities, and systematic failure modes across speaking rates and acoustic environments.
Problem

Research questions and friction points this paper is trying to address.

Indian Languages
Automatic Speech Recognition
Evaluation Benchmark
Linguistic Variations
Demographic Diversity
Innovation

Methods, ideas, or system contributions that make the work stand out.

demographically diverse on-field recordings
in-the-wild audio
lattice of variations framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.