FFASR: Benchmarking Far-Field Automatic Speech Recognition using High-Fidelity Simulated RIRs

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of standardized far-field automatic speech recognition (ASR) benchmarks and the inadequacy of existing evaluations in assessing reverberation and moving-speaker effects. To this end, we construct a far-field ASR benchmark and leaderboard comprising 15,637 utterances. Methodologically, hybrid wave-geometric acoustic simulation is employed to generate high-fidelity room impulse responses across fourteen rooms, covering nine single-variable acoustic conditions. Additionally, a newly recorded test set designed to prevent data contamination validates the simulation's effectiveness. Results demonstrate that word error rates (WER) reach 41.3% under low signal-to-noise ratio static conditions, with moving speakers introducing further performance degradation. The average discrepancy between simulated and measured WER is merely 1.7 percentage points, confirming that high-fidelity simulation serves as a scalable proxy for real-world far-field evaluation.
📝 Abstract
Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emphasize close-microphone speech. We present FFASR, a held-out corpus of 15,637 utterances and an open leaderboard spanning nine conditions, each varying a single acoustic factor: anechoic near-field speech, a measured-versus-simulated office-lab pair, static far-field mixtures at high/mid/low signal-to-noise ratio(SNR), and moving-talker variants at matched SNR. Dry speech from 15 talkers is convolved with hybrid wave/geometrical-acoustics room impulse responses from 14 furnished rooms; because the speech is newly recorded and the test waveforms are never released, the corpus resists training-data contamination. Across contemporary systems, mean word error rate (WER) rises from 4.4% near-field to 41.3% in the static low-SNR condition; a moving talker adds a small but consistent penalty at matched SNR; and on the office-lab pair, measured and simulated WER agree to within about 1.7 pp on average. These results support high-fidelity simulation as a scalable proxy for measured far-field evaluation under the conditions we test.
Problem

Research questions and friction points this paper is trying to address.

far-field automatic speech recognition
benchmarking
reverberation
noise
talker motion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Far-Field ASR
Room Impulse Response
High-Fidelity Simulation
Benchmark
Moving Talker
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shivam Saini
Treble Technologies, Reykjavík, Iceland
Eric Bezzam
Eric Bezzam
EPFL
Signal processingmachine learningcomputational imagingaudio
Georg Götz
Georg Götz
Treble Technologies
A
Alessia Milo
Treble Technologies, Reykjavík, Iceland
S
Steinar Guðjónsson
Treble Technologies, Reykjavík, Iceland
K
Konstantinos Gkanos
Treble Technologies, Reykjavík, Iceland
F
Finnur Pind
Treble Technologies, Reykjavík, Iceland
Daniel Gert Nielsen
Daniel Gert Nielsen
Ph.D student Acoustics
OptimizationNumerical ModellingVibro-acoustics