annotate asr hallucinations

Designs and builds human‑annotated corpora of ASR hallucinations by labeling hallucinated spans in speech transcripts, assigning severity and other metadata, linking labels to the corresponding source audio segments, and consolidating annotations collected across multiple ASR models. Produces annotation guidelines and dataset artifacts that enable analysis of hallucination types, frequencies, and model comparisons.

annotateasrhallucinations

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the susceptibility of end-to-end automatic speech recognition (ASR) systems to hallucination in natural speech, noting that existing mitigation approaches are predominantly evaluated on synthetic or artificially degraded audio, lacking benchmarks from real-world scenarios. To bridge this gap, the authors introduce HALAS, the first human-annotated dataset for ASR hallucination based on authentic, unprocessed earnings call recordings, covering seven state-of-the-art models and providing segment-level labels to analyze hallucination patterns and severity. Using character- and semantic-level metrics alongside ROC-AUC and F1 scores, the study reveals substantial overlap in hallucinated terms across models and significant hallucination even in transcriptions with low word error rates. Experiments on HALAS show that the proposed detection metrics achieve 81% ROC-AUC, while the best existing method attains only 53.1% F1, highlighting the limitations of current approaches in realistic settings.

ASR hallucinationsautomatic speech recognitionhallucination detection

Hallucination Benchmark for Speech Foundation Models

Oct 18, 2025
AK
Alkis Koudounas
🏛️ Politecnico di Torino | Kore University of Enna | Amazon AGI | Università degli Studi di Palermo

Hallucinations in automatic speech recognition (ASR)—i.e., fluent yet semantically unrelated transcriptions—pose severe risks in high-stakes domains (e.g., healthcare, law) due to their stealthy nature; conventional word error rate (WER)-based evaluation fails to detect them. Method: We propose SHALLOW, the first systematic hallucination assessment framework for ASR, modeling hallucinatory characteristics across four orthogonal dimensions: lexical, phonetic, morphological, and semantic, yielding an interpretable, multi-axis quantitative metric. Contribution/Results: SHALLOW enables fine-grained hallucination classification and measurement, exposes WER’s breakdown under high acoustic noise, and supports precise diagnostic analysis of model behavior. Experiments show SHALLOW strongly correlates with WER at low error rates but remains discriminative across models under high WER—significantly enhancing detection sensitivity and analytical depth for non-faithful outputs.

Creating metrics to evaluate lexical phonetic morphological semantic errorsDeveloping benchmark to detect ASR hallucinations in speech modelsProviding fine-grained error analysis beyond word error rate limitations

Language models should be subject to repeatable, open, domain-contextualized hallucination benchmarking

May 22, 2025
JD
Justin D. Norman
🏛️ University of California, Berkeley

Current evaluation frameworks for language model hallucinations lack systematicity and reproducibility, often脱离 real-world domain contexts and suffering from low validity. Method: We propose the first domain-contextualized, open, and reproducible hallucination evaluation framework, featuring: (1) an original hallucination taxonomy; (2) an expert-collaborative annotation protocol, empirically demonstrating that early expert involvement critically enhances evaluation validity; and (3) context-sensitive benchmark tasks with rigorous validity verification procedures. Results: Experiments reveal that non-expert–driven evaluations frequently distort metric outcomes. Our framework substantially improves reliability, construct validity, and cross-domain generalizability of hallucination assessment. It provides both theoretical foundations and a practical paradigm for developing high-trust hallucination benchmarks.

Address validity issues in hallucination metrics without expert inputEvaluate models with repeatable open contextualized benchmarkingMeasure prevalence of language model hallucination comprehensively

Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Sep 03, 2023
YZ
Yue Zhang
🏛️ Soochow University | Zhejiang University | Tencent AI Lab | Renmin University of China | Nanyang Technological University

Hallucination—i.e., generation inconsistent with input, context, or factual knowledge—severely undermines the reliability of large language models (LLMs) in real-world applications. To address this, we systematically analyze hallucination causes and propose, for the first time, a multidimensional taxonomy tailored to LLMs, encompassing input consistency, contextual coherence, and factual alignment, alongside a unified cross-task evaluation benchmark. Synthesizing insights from over 120 studies, we integrate and comparatively assess major mitigation strategies—including confidence calibration, retrieval-augmented generation (RAG), self-verification, contrastive decoding, and knowledge graph alignment—revealing their empirical limitations in open-domain question answering and long-horizon reasoning. Our work establishes a novel paradigm of synergistic governance via interpretable intervention and knowledge enhancement, offering both a theoretical framework and practical guidelines for developing trustworthy LLMs.

Addressing hallucination issues in large language modelsDetecting and mitigating content divergence from user inputImproving reliability of LLMs in real-world applications

Latest Papers

What's happening recently
View more

This work addresses the critical issue of fluent yet audio-irrelevant hallucinations generated by automatic speech recognition (ASR) models such as Whisper, which undermine system reliability. The study presents the first systematic comparison of three reference-free hallucination detection approaches: text-based metrics, large language model prompting strategies, and probing internal states of the Whisper decoder. Furthermore, it introduces a lightweight late-fusion meta-classifier that effectively integrates signals from multiple sources. Experimental results reveal that intermediate-layer representations within the Whisper decoder alone exhibit strong discriminative power for hallucination detection. Moreover, the proposed meta-classifier, which fuses textual and internal decoder features, achieves the best trade-off between precision and recall, significantly outperforming existing methods.

ASR hallucinationautomatic speech recognitionhallucination detection

研究解决了ASR系统产生无关文本的问题,通过分析两个Conformer-Large模型在不同条件下的表现,发现最终编码阶段是关键,其失败导致输出失去音频基础。

ASRgrounding failurehallucination

This study addresses the challenges of localizing hallucinations and retrieving supporting evidence in large language models by proposing a novel joint modeling framework for hallucination detection and input evidence alignment. Integrating encoder masked prediction, confidence estimation, and token-level alignment mechanisms, the approach enables fine-grained hallucination identification with interpretable traceability. This method overcomes limitations of existing detection techniques by precisely pinpointing hallucinated segments while associating them with high-quality input evidence. Human evaluation confirms the reliability of the alignment results, demonstrating significant improvements in the trustworthiness and transparency of generated content. Collectively, this work establishes a new paradigm for mitigating hallucinations through verifiable evidence grounding.

Conditional Text GenerationHallucination Span DetectionInput-Side Evidence Alignment

This work addresses the limitations of existing hallucination evaluation methods for vision-language models, which predominantly rely on model-generated samples and suffer from poor timeliness and insufficient controllability. To overcome these issues, the authors construct a dataset comprising 1,600 human-written, multilingual hallucination samples, complemented by fine-grained, span-level annotations. Through systematic comparative analysis, they demonstrate for the first time that human-authored samples significantly outperform model-generated ones in terms of distributional similarity, annotation consistency, and content controllability. These advantages enable more stable and generalizable assessment of model hallucination detection capabilities, offering a reliable alternative for benchmarking hallucinations in vision-language systems.

fine-grained annotationhallucination benchmarkinghuman-written samples

This work addresses the challenge that existing audio-visual large language models struggle to accurately align spoken content with visual signals, often generating semantic or temporal hallucinations. To this end, the paper introduces SVHalluc, the first benchmark specifically designed to evaluate speech–vision hallucination, systematically assessing cross-modal understanding along two dimensions: semantic consistency and temporal alignment. The benchmark employs a multi-task framework and includes comprehensive experiments covering both open-source and proprietary state-of-the-art models. Results reveal that current open-source models perform near random chance, while Gemini 2.5 Pro demonstrates significant superiority, highlighting a critical gap in existing approaches for speech–vision alignment and establishing a much-needed evaluation standard in this emerging domain.

audio-visual LLMscross-modality alignmentsemantic grounding

Hot Scholars

CJ

Chathuri Jayaweera

University of Florida
Natural Language InferenceAmbiguity DetectionCommonsense Reasoning
SY

Sangpil Youm

Ph.D Student, University of Florida
Natural Language ProcessingArtificial IntelligenceNetwork Science
BD

Bonnie Dorr

University of Florida, IHMC, UMD
Artificial IntelligenceNatural Language ProcessingMachine Translation