Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results

๐Ÿ“… 2026-07-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Conventional wisdom holds that human auditory recognition performance represents the upper bound for automatic speech recognition (ASR) systems, yet this assumption lacks systematic validation in diverse speech contextsโ€”such as those involving children, older adults, and regional accents. This study presents the first systematic comparison between native Dutch listeners and state-of-the-art ASR systems (e.g., Google Telephony) on authentic, diverse Dutch speech data, examining the effects of speaker age, regional accent, and utterance length. Results reveal that ASR performance is generally on par with human listeners and even surpasses them under certain conditions, with system accuracy highly sensitive to test set composition. These findings challenge the long-standing belief that humans consistently outperform machines in speech recognition and underscore the critical need to enhance ASR robustness to speaker age and accent variation.
๐Ÿ“ Abstract
Humans are often considered to be the best listeners and seen as the upper-bound performance of automatic speech recognition (ASR) systems. We present a preliminary comparison of the performances of state-of-the-art ASR systems and Dutch native listeners on the recognition of "diverse" speech, specifically Dutch child and older adults' speech and Flemish. Google Telephony outperformed the other ASR systems. Importantly, the ASR systems showed similar performance to the listeners, and in specific cases even outperformed them. Slight performance differences between the listeners and ASR systems were found related to speaker's age and regional accents and utterance length. Future research should focus on making ASR systems more robust to acoustic variability related to aging and regional accents. A comparison of ASR recognition performances on the test stimuli and the full Jasmin-CGN test sets showed the influence of the specific test sets on the conclusions regarding benchmarking human and ASR performance.
Problem

Research questions and friction points this paper is trying to address.

automatic speech recognition
diverse speech
speaker age
regional accents
human benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

automatic speech recognition
diverse speech
human benchmarking
regional accents
speaker age
๐Ÿ”Ž Similar Papers