Score
Design and conduct analyses of automatic speech recognition (ASR) transcription errors that categorize and label errors by phonetic cause, build error taxonomies and confusion matrices, and quantify how phonetic substitutions affect downstream processing. Use linear mixed‑effects models to estimate fixed and random effects on error rates, control for speaker/item variability, and produce statistical summaries that guide data augmentation and system improvements.
This study addresses the robustness evaluation of automatic speech recognition (ASR) systems against spontaneous speech errors. We introduce SFUSED—the first English spontaneous speech error corpus with multi-level linguistic annotations (word- and syllable-level error localization, context sensitivity, correction patterns), comprising 5,300 utterances—and employ its structured error annotation schema for ASR diagnostics, conducting a systematic evaluation of WhisperX. Methodologically, we integrate degraded word detection, error type classification, and correction behavior modeling to enable fine-grained error attribution. Results demonstrate that SFUSED effectively exposes ASR vulnerabilities in realistic conversational contexts: WhisperX exhibits strong robustness against repetitions and fillers but shows significant limitations in handling phonological illusions and context-dependent corrections. This work establishes a reproducible benchmark framework and diagnostic paradigm for ASR robustness assessment.
Researchers have demonstrated that Automatic Speech Recognition (ASR) systems perform differently across demographic groups. In this work, we examined how subtitle errors affect evaluations of speakers and their content using a preregistered online experiment (N=207, U.S.-based crowdworkers). Participants watched speakers with various accents deliver a talk in which the subtitles were accurate or error-prone. Our results indicate that error-prone subtitles consistently reduce both speaker and content evaluations for all speakers. We did not see disparate impact between the accent groups, controlling for subtitle quality. Taken together, though, the findings of this short paper imply that speakers with accents for which ASR systems perform poorly are likely to be further penalized by viewers with lower evaluations.
Automated speech classification errors systematically bias statistical inference in child language development research—particularly estimates of sibling effects and input–output associations. This paper quantifies, for the first time, the magnitude of bias induced by prevalent classifiers (e.g., LENA, ACLEW) on regression effect sizes (e.g., correlations, coefficients), revealing that their misclassification attenuates the estimated negative effect of siblings on adult language input by 20–80%. To address this, we propose a Bayesian calibration framework that models the classifier’s confusion matrix using manually verified annotations, enabling unbiased estimation of target effects. Our method substantially reduces estimation bias and establishes a generalizable error-analysis paradigm for event-detection classifiers. By rigorously characterizing and correcting classifier-induced distortion, this work advances the methodological rigor of automated speech annotation in developmental science.
This study investigates the effectiveness and robustness of automatic speech recognition (ASR) transcripts for speaker attribution, particularly examining the impact of transcription errors. Contrary to the conventional assumption that ASR errors degrade performance, experiments reveal that word-level errors do not significantly impair attribution accuracy—and may even introduce speaker-discriminative linguistic cues; in certain settings, ASR-generated transcripts yield superior attribution compared to human transcriptions. The work systematically evaluates how ASR system characteristics—including recognition accuracy, vocabulary coverage, and contextual modeling capability—affect attribution performance, corroborating findings through linguistic pattern analysis and speaker特征 modeling. Results demonstrate that speaker attribution based on ASR output is highly resilient and practically viable, challenging the prevailing assumption that high-fidelity transcription is a prerequisite. This establishes a novel paradigm for speaker identification in audio-deprived scenarios, where only ASR transcripts are available.
This study systematically evaluates racial bias in four major commercial automatic speech recognition (ASR) systems. Using the Pacific Northwest English Corpus—featuring speakers from African American, White, Chicano, and Yakama communities—the authors quantify cross-ethnic transcription disparities via a novel phoneme error rate (PER) metric integrated with sociophonetic annotations. Results reveal significantly higher PER for African American speakers across all systems; critically, all models substantially underrepresent dialectal phenomena such as low-vowel mergers, confirming that inadequate acoustic modeling of sociophonetic variation constitutes the primary source of bias. The study introduces an analytical framework linking PER to fine-grained sociophonetic features, identifying vowel quality variation as a key determinant of performance disparity. These findings underscore the necessity of incorporating dialectal diversity into ASR training and evaluation to advance fairness and robustness in speech technology.
This work addresses the challenge of atypical speech recognition, where conventional evaluation practices relying on a single transcription standard conflate verbatim fidelity with semantic intent. To resolve this ambiguity, the study introduces the first dual-reference benchmark framework specifically designed for atypical speech such as stuttering, incorporating both verbatim transcripts (preserving repetitions and prolongations) and intent-based transcripts (normalized and redundancy-free). The authors systematically evaluate eleven state-of-the-art ASR models spanning encoder-decoder, CTC, and Transducer architectures under both reference types. Results reveal substantial discrepancies in model performance and ranking depending on the chosen transcription standard, demonstrating that the selection of evaluation criteria critically influences the real-world deployment of atypical ASR systems and providing empirical guidance for context-appropriate model selection.
This study addresses the limitations of traditional automatic speech recognition (ASR) evaluation, which relies heavily on word error rate (WER) and fails to capture the grammatical and semantic characteristics of transcription errors. To overcome this, the authors propose two novel metrics: Part-of-Speech Error Rate (POSER) and Embedding Error Rate (EmbER), which quantify ASR output quality from the perspectives of grammatical correctness and semantic fidelity, respectively. By integrating language model rescoring, part-of-speech tagging, and semantic distance computation based on word embeddings, they construct a multidimensional qualitative evaluation framework. Experimental results demonstrate that these new metrics effectively reveal the contribution of language models to improving linguistic quality in transcriptions, thereby compensating for WER’s insufficiency in linguistic analysis.
Traditional automatic speech recognition evaluation metrics, such as word error rate (WER) and character error rate (CER), fail to capture human perception of errors and neglect linguistic and semantic influences. This work proposes a novel paradigm that embeds any perception-oriented evaluation metric into the minimum edit distance (minED) framework to produce an intuitively interpretable equivalent error rate. For the first time, this approach translates human perceptual modeling into a comprehensible error rate format, enabling quantification of error severity from the perspective of human understanding. The resulting metric not only aligns closely with human judgments but also effectively identifies recognition errors that critically impact semantic comprehension.
This study investigates the propagation of automatic speech recognition (ASR) errors in Korean spoken question-answering systems that employ an ASR–large language model (LLM) cascade, and the resulting semantic failures. Through ASR error analysis, semantic failure evaluation, and comparative experiments with end-to-end audio-language models, the authors demonstrate that even single-character ASR errors can cause complete downstream QA failure. They find that information loss during ASR is the primary driver of performance degradation, and LLMs of varying capabilities exhibit similar sensitivity to such errors. The results indicate that end-to-end models directly processing audio inputs significantly outperform conventional cascaded architectures in noisy conditions, effectively mitigating semantic information loss caused by transcription errors.
This study addresses the limited performance of automatic speech recognition (ASR) for low-resource languages and the unclear efficacy of large language models (LLMs) in generative error correction (GER), particularly under data contamination concerns. For the first time, the authors establish a rigorously decontaminated offline evaluation benchmark for West Frisian by leveraging both public and private corpora, enabling a systematic assessment of LLM-based GER on ASR outputs. Experimental results demonstrate that GER significantly improves ASR performance across most configurations, with GPT-5.1 even surpassing the oracle word error rate—a finding that underscores the robustness and effectiveness of LLMs in correcting ASR errors in low-resource settings.