train and evaluate asr

Designs and trains automatic speech recognition (ASR) systems and associated training pipelines—including dataset partitioning to avoid speaker leakage, adaptation and fine‑tuning workflows, and speaker‑independent or personalized variants—and implements mechanisms for detecting and characterizing hallucinations. Builds evaluation and benchmarking suites that analyze ASR errors and robustness across speakers and conditions, compute recognition and perceptual metrics (e.g., WER/CER and speech quality measures), run hallucination‑detector evaluations, and perform human–machine listening comparisons to assess generalization and system behavior.

trainandevaluateasr

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$191K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Hallucination Benchmark for Speech Foundation Models

Oct 18, 2025
AK
Alkis Koudounas
🏛️ Politecnico di Torino | Kore University of Enna | Amazon AGI | Università degli Studi di Palermo

Hallucinations in automatic speech recognition (ASR)—i.e., fluent yet semantically unrelated transcriptions—pose severe risks in high-stakes domains (e.g., healthcare, law) due to their stealthy nature; conventional word error rate (WER)-based evaluation fails to detect them. Method: We propose SHALLOW, the first systematic hallucination assessment framework for ASR, modeling hallucinatory characteristics across four orthogonal dimensions: lexical, phonetic, morphological, and semantic, yielding an interpretable, multi-axis quantitative metric. Contribution/Results: SHALLOW enables fine-grained hallucination classification and measurement, exposes WER’s breakdown under high acoustic noise, and supports precise diagnostic analysis of model behavior. Experiments show SHALLOW strongly correlates with WER at low error rates but remains discriminative across models under high WER—significantly enhancing detection sensitivity and analytical depth for non-faithful outputs.

Creating metrics to evaluate lexical phonetic morphological semantic errorsDeveloping benchmark to detect ASR hallucinations in speech modelsProviding fine-grained error analysis beyond word error rate limitations

Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio

Jan 20, 2025
MB
Mateusz Barański
🏛️ AGH University of Krakow

This study identifies a critical hallucination vulnerability in the Whisper automatic speech recognition (ASR) model when exposed to non-speech audio—such as environmental sounds. Systematically injecting diverse non-speech adversarial signals, we characterize, for the first time, Whisper’s hallucination sensitivity patterns across noise types. To address this without model retraining, we propose the Bag-of-Hallucinations (BoH), a lightweight post-processing framework that statistically models and identifies high-frequency hallucinated tokens, enabling transcription purification. BoH integrates ASR robustness analysis, adversarial audio injection, and token-level hallucination modeling. Experiments across multiple real-world noise conditions demonstrate that BoH significantly reduces word error rate (WER), effectively suppresses hallucinated outputs, and enhances transcription reliability and safety. Our approach establishes a novel paradigm for trustworthy deployment of large-scale ASR models in open-domain, non-ideal acoustic environments.

Non-speech SoundsRecognition ErrorsWhisper System

This work addresses hallucination in speech foundation models for automatic speech recognition (ASR)—i.e., generation of text severely inconsistent with the input audio—a critical safety hazard in high-stakes domains such as healthcare and law. Conventional metrics (e.g., WER, CER) fail to capture hallucination meaningfully; thus, we propose the **Hallucination Error Rate (HER)**, the first formally defined and quantifiable metric for this phenomenon. Systematic evaluation across 20 state-of-the-art ASR models reveals HER’s strong correlation with input distribution shift (α = 0.91), while low WER often masks substantial hallucination risk. Further validation via distribution modeling, synthetic noise robustness testing, and adversarial perturbation analysis demonstrates that HER more reliably reflects true hallucination propensity than standard metrics. Our findings establish HER as a discriminative, safety-aware evaluation paradigm for ASR in high-risk applications.

Address hallucination in ASR modelsAssess ASR performance in high-stakes domainsIntroduce HER to quantify hallucinations

This work addresses the susceptibility of end-to-end automatic speech recognition (ASR) systems to hallucination in natural speech, noting that existing mitigation approaches are predominantly evaluated on synthetic or artificially degraded audio, lacking benchmarks from real-world scenarios. To bridge this gap, the authors introduce HALAS, the first human-annotated dataset for ASR hallucination based on authentic, unprocessed earnings call recordings, covering seven state-of-the-art models and providing segment-level labels to analyze hallucination patterns and severity. Using character- and semantic-level metrics alongside ROC-AUC and F1 scores, the study reveals substantial overlap in hallucinated terms across models and significant hallucination even in transcriptions with low word error rates. Experiments on HALAS show that the proposed detection metrics achieve 81% ROC-AUC, while the best existing method attains only 53.1% F1, highlighting the limitations of current approaches in realistic settings.

ASR hallucinationsautomatic speech recognitionhallucination detection

This work addresses the critical issue of fluent yet audio-irrelevant hallucinations generated by automatic speech recognition (ASR) models such as Whisper, which undermine system reliability. The study presents the first systematic comparison of three reference-free hallucination detection approaches: text-based metrics, large language model prompting strategies, and probing internal states of the Whisper decoder. Furthermore, it introduces a lightweight late-fusion meta-classifier that effectively integrates signals from multiple sources. Experimental results reveal that intermediate-layer representations within the Whisper decoder alone exhibit strong discriminative power for hallucination detection. Moreover, the proposed meta-classifier, which fuses textual and internal decoder features, achieves the best trade-off between precision and recall, significantly outperforming existing methods.

ASR hallucinationautomatic speech recognitionhallucination detection

Latest Papers

What's happening recently
View more

This work addresses the tendency of Whisper speech recognition models to generate fluent yet erroneous “hallucinated” transcriptions on non-speech audio. The study reveals, for the first time, that hallucination signals are linearly separable within the encoder’s hidden representations and predominantly localized in a sparse subset of features. Building on this insight, the authors propose a fine-tuning-free latent-space intervention strategy that leverages both raw activations and latent representations from a sparse autoencoder (SAE) to detect and suppress hallucinations. Experimental results demonstrate that this approach reduces hallucination rates on non-speech test sets from 72.63% to 14.11% for Whisper-small and from 86.88% to 27.33% for Whisper-large-v3, with only a marginal increase in word error rate (WER), achieving performance comparable to fine-tuned baselines.

ASRhallucinationnon-speech audio

Hot Scholars

SW

Shinji Watanabe

Carnegie Mellon University
Speech recognitionSpeech processingSpeech enhancementSpeech translation
HY

Hung-yi Lee

National Taiwan University
deep learningspoken language understandingspeech processing
ZN

Zhikang Niu

Shanghai Jiao Tong University
Speech Synthesis
LX

Lei Xie

Northwestern Polytechnical University
speech processingspeech recognitionspeech synthesismultimedia
XL

Xunying Liu

Chinese University of Hong Kong
Speech and Language ProcessingMachine Learning