Score
Designs and trains automatic speech recognition (ASR) systems and associated training pipelines—including dataset partitioning to avoid speaker leakage, adaptation and fine‑tuning workflows, and speaker‑independent or personalized variants—and implements mechanisms for detecting and characterizing hallucinations. Builds evaluation and benchmarking suites that analyze ASR errors and robustness across speakers and conditions, compute recognition and perceptual metrics (e.g., WER/CER and speech quality measures), run hallucination‑detector evaluations, and perform human–machine listening comparisons to assess generalization and system behavior.
Hallucinations in automatic speech recognition (ASR)—i.e., fluent yet semantically unrelated transcriptions—pose severe risks in high-stakes domains (e.g., healthcare, law) due to their stealthy nature; conventional word error rate (WER)-based evaluation fails to detect them. Method: We propose SHALLOW, the first systematic hallucination assessment framework for ASR, modeling hallucinatory characteristics across four orthogonal dimensions: lexical, phonetic, morphological, and semantic, yielding an interpretable, multi-axis quantitative metric. Contribution/Results: SHALLOW enables fine-grained hallucination classification and measurement, exposes WER’s breakdown under high acoustic noise, and supports precise diagnostic analysis of model behavior. Experiments show SHALLOW strongly correlates with WER at low error rates but remains discriminative across models under high WER—significantly enhancing detection sensitivity and analytical depth for non-faithful outputs.
This study identifies a critical hallucination vulnerability in the Whisper automatic speech recognition (ASR) model when exposed to non-speech audio—such as environmental sounds. Systematically injecting diverse non-speech adversarial signals, we characterize, for the first time, Whisper’s hallucination sensitivity patterns across noise types. To address this without model retraining, we propose the Bag-of-Hallucinations (BoH), a lightweight post-processing framework that statistically models and identifies high-frequency hallucinated tokens, enabling transcription purification. BoH integrates ASR robustness analysis, adversarial audio injection, and token-level hallucination modeling. Experiments across multiple real-world noise conditions demonstrate that BoH significantly reduces word error rate (WER), effectively suppresses hallucinated outputs, and enhances transcription reliability and safety. Our approach establishes a novel paradigm for trustworthy deployment of large-scale ASR models in open-domain, non-ideal acoustic environments.
This work addresses hallucination in speech foundation models for automatic speech recognition (ASR)—i.e., generation of text severely inconsistent with the input audio—a critical safety hazard in high-stakes domains such as healthcare and law. Conventional metrics (e.g., WER, CER) fail to capture hallucination meaningfully; thus, we propose the **Hallucination Error Rate (HER)**, the first formally defined and quantifiable metric for this phenomenon. Systematic evaluation across 20 state-of-the-art ASR models reveals HER’s strong correlation with input distribution shift (α = 0.91), while low WER often masks substantial hallucination risk. Further validation via distribution modeling, synthetic noise robustness testing, and adversarial perturbation analysis demonstrates that HER more reliably reflects true hallucination propensity than standard metrics. Our findings establish HER as a discriminative, safety-aware evaluation paradigm for ASR in high-risk applications.
This work addresses the susceptibility of end-to-end automatic speech recognition (ASR) systems to hallucination in natural speech, noting that existing mitigation approaches are predominantly evaluated on synthetic or artificially degraded audio, lacking benchmarks from real-world scenarios. To bridge this gap, the authors introduce HALAS, the first human-annotated dataset for ASR hallucination based on authentic, unprocessed earnings call recordings, covering seven state-of-the-art models and providing segment-level labels to analyze hallucination patterns and severity. Using character- and semantic-level metrics alongside ROC-AUC and F1 scores, the study reveals substantial overlap in hallucinated terms across models and significant hallucination even in transcriptions with low word error rates. Experiments on HALAS show that the proposed detection metrics achieve 81% ROC-AUC, while the best existing method attains only 53.1% F1, highlighting the limitations of current approaches in realistic settings.
This work addresses the critical issue of fluent yet audio-irrelevant hallucinations generated by automatic speech recognition (ASR) models such as Whisper, which undermine system reliability. The study presents the first systematic comparison of three reference-free hallucination detection approaches: text-based metrics, large language model prompting strategies, and probing internal states of the Whisper decoder. Furthermore, it introduces a lightweight late-fusion meta-classifier that effectively integrates signals from multiple sources. Experimental results reveal that intermediate-layer representations within the Whisper decoder alone exhibit strong discriminative power for hallucination detection. Moreover, the proposed meta-classifier, which fuses textual and internal decoder features, achieves the best trade-off between precision and recall, significantly outperforming existing methods.
This work addresses the tendency of Whisper speech recognition models to generate fluent yet erroneous “hallucinated” transcriptions on non-speech audio. The study reveals, for the first time, that hallucination signals are linearly separable within the encoder’s hidden representations and predominantly localized in a sparse subset of features. Building on this insight, the authors propose a fine-tuning-free latent-space intervention strategy that leverages both raw activations and latent representations from a sparse autoencoder (SAE) to detect and suppress hallucinations. Experimental results demonstrate that this approach reduces hallucination rates on non-speech test sets from 72.63% to 14.11% for Whisper-small and from 86.88% to 27.33% for Whisper-large-v3, with only a marginal increase in word error rate (WER), achieving performance comparable to fine-tuned baselines.