Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models

πŸ“… 2025-02-18
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses hallucination in speech foundation models for automatic speech recognition (ASR)β€”i.e., generation of text severely inconsistent with the input audioβ€”a critical safety hazard in high-stakes domains such as healthcare and law. Conventional metrics (e.g., WER, CER) fail to capture hallucination meaningfully; thus, we propose the **Hallucination Error Rate (HER)**, the first formally defined and quantifiable metric for this phenomenon. Systematic evaluation across 20 state-of-the-art ASR models reveals HER’s strong correlation with input distribution shift (Ξ± = 0.91), while low WER often masks substantial hallucination risk. Further validation via distribution modeling, synthetic noise robustness testing, and adversarial perturbation analysis demonstrates that HER more reliably reflects true hallucination propensity than standard metrics. Our findings establish HER as a discriminative, safety-aware evaluation paradigm for ASR in high-risk applications.

Technology Category

Natural Language Processing: Safety and RobustnessMachine Learning: Evaluation and AnalysisCognitive Modeling & Cognitive Systems: Other Foundations of Cognitive Modeling & Systems

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSecurity and Privacy: Large-scale security measurementsUser Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating success
πŸ“ Abstract
Speech foundation models trained at a massive scale, both in terms of model and data size, result in robust systems capable of performing multiple speech tasks, including automatic speech recognition (ASR). These models transcend language and domain barriers, yet effectively measuring their performance remains a challenge. Traditional metrics like word error rate (WER) and character error rate (CER) are commonly used to evaluate ASR performance but often fail to reflect transcription quality in critical contexts, particularly when detecting fabricated outputs. This phenomenon, known as hallucination, is especially concerning in high-stakes domains such as healthcare, legal, and aviation, where errors can have severe consequences. In our work, we address this gap by investigating hallucination in ASR models. We examine how factors such as distribution shifts, model size, and model architecture influence the hallucination error rate (HER), a metric we introduce to quantify hallucinations. Our analysis of 20 ASR models reveals uminsights~key insights: (1) High WERs can mask low hallucination rates, while low WERs may conceal dangerous hallucinations. (2) Synthetic noise, both adversarial and common perturbations like white noise, pitch shift, and time stretching, increase HER. (3) Distribution shift correlates strongly with HER ($alpha = 0.91$). Our findings highlight the importance of incorporating HER alongside traditional metrics like WER to better assess ASR model performance, particularly in high-stakes domains.
Problem

Research questions and friction points this paper is trying to address.

Address hallucination in ASR models
Introduce HER to quantify hallucinations
Assess ASR performance in high-stakes domains
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduced HER metric for hallucinations.
Analyzed distribution shift impact on HER.
Highlighted synthetic noise effects on HER.
πŸ”Ž Similar Papers
No similar papers found.