🤖 AI Summary
This study addresses the absence of standardized far-field automatic speech recognition (ASR) benchmarks and the inadequacy of existing evaluations in assessing reverberation and moving-speaker effects. To this end, we construct a far-field ASR benchmark and leaderboard comprising 15,637 utterances. Methodologically, hybrid wave-geometric acoustic simulation is employed to generate high-fidelity room impulse responses across fourteen rooms, covering nine single-variable acoustic conditions. Additionally, a newly recorded test set designed to prevent data contamination validates the simulation's effectiveness. Results demonstrate that word error rates (WER) reach 41.3% under low signal-to-noise ratio static conditions, with moving speakers introducing further performance degradation. The average discrepancy between simulated and measured WER is merely 1.7 percentage points, confirming that high-fidelity simulation serves as a scalable proxy for real-world far-field evaluation.
📝 Abstract
Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emphasize close-microphone speech. We present FFASR, a held-out corpus of 15,637 utterances and an open leaderboard spanning nine conditions, each varying a single acoustic factor: anechoic near-field speech, a measured-versus-simulated office-lab pair, static far-field mixtures at high/mid/low signal-to-noise ratio(SNR), and moving-talker variants at matched SNR. Dry speech from 15 talkers is convolved with hybrid wave/geometrical-acoustics room impulse responses from 14 furnished rooms; because the speech is newly recorded and the test waveforms are never released, the corpus resists training-data contamination. Across contemporary systems, mean word error rate (WER) rises from 4.4% near-field to 41.3% in the static low-SNR condition; a moving talker adds a small but consistent penalty at matched SNR; and on the office-lab pair, measured and simulated WER agree to within about 1.7 pp on average. These results support high-fidelity simulation as a scalable proxy for measured far-field evaluation under the conditions we test.