Score
Designs and builds human‑annotated corpora of ASR hallucinations by labeling hallucinated spans in speech transcripts, assigning severity and other metadata, linking labels to the corresponding source audio segments, and consolidating annotations collected across multiple ASR models. Produces annotation guidelines and dataset artifacts that enable analysis of hallucination types, frequencies, and model comparisons.
This work addresses the susceptibility of end-to-end automatic speech recognition (ASR) systems to hallucination in natural speech, noting that existing mitigation approaches are predominantly evaluated on synthetic or artificially degraded audio, lacking benchmarks from real-world scenarios. To bridge this gap, the authors introduce HALAS, the first human-annotated dataset for ASR hallucination based on authentic, unprocessed earnings call recordings, covering seven state-of-the-art models and providing segment-level labels to analyze hallucination patterns and severity. Using character- and semantic-level metrics alongside ROC-AUC and F1 scores, the study reveals substantial overlap in hallucinated terms across models and significant hallucination even in transcriptions with low word error rates. Experiments on HALAS show that the proposed detection metrics achieve 81% ROC-AUC, while the best existing method attains only 53.1% F1, highlighting the limitations of current approaches in realistic settings.
Hallucinations in automatic speech recognition (ASR)—i.e., fluent yet semantically unrelated transcriptions—pose severe risks in high-stakes domains (e.g., healthcare, law) due to their stealthy nature; conventional word error rate (WER)-based evaluation fails to detect them. Method: We propose SHALLOW, the first systematic hallucination assessment framework for ASR, modeling hallucinatory characteristics across four orthogonal dimensions: lexical, phonetic, morphological, and semantic, yielding an interpretable, multi-axis quantitative metric. Contribution/Results: SHALLOW enables fine-grained hallucination classification and measurement, exposes WER’s breakdown under high acoustic noise, and supports precise diagnostic analysis of model behavior. Experiments show SHALLOW strongly correlates with WER at low error rates but remains discriminative across models under high WER—significantly enhancing detection sensitivity and analytical depth for non-faithful outputs.
研究针对低资源语言的口语幻觉检测问题,构建了包含英、俄、哈三种语言的多语种数据集,并通过文本和直接音频处理方法进行评估。
Current evaluation frameworks for language model hallucinations lack systematicity and reproducibility, often脱离 real-world domain contexts and suffering from low validity. Method: We propose the first domain-contextualized, open, and reproducible hallucination evaluation framework, featuring: (1) an original hallucination taxonomy; (2) an expert-collaborative annotation protocol, empirically demonstrating that early expert involvement critically enhances evaluation validity; and (3) context-sensitive benchmark tasks with rigorous validity verification procedures. Results: Experiments reveal that non-expert–driven evaluations frequently distort metric outcomes. Our framework substantially improves reliability, construct validity, and cross-domain generalizability of hallucination assessment. It provides both theoretical foundations and a practical paradigm for developing high-trust hallucination benchmarks.
Hallucination—i.e., generation inconsistent with input, context, or factual knowledge—severely undermines the reliability of large language models (LLMs) in real-world applications. To address this, we systematically analyze hallucination causes and propose, for the first time, a multidimensional taxonomy tailored to LLMs, encompassing input consistency, contextual coherence, and factual alignment, alongside a unified cross-task evaluation benchmark. Synthesizing insights from over 120 studies, we integrate and comparatively assess major mitigation strategies—including confidence calibration, retrieval-augmented generation (RAG), self-verification, contrastive decoding, and knowledge graph alignment—revealing their empirical limitations in open-domain question answering and long-horizon reasoning. Our work establishes a novel paradigm of synergistic governance via interpretable intervention and knowledge enhancement, offering both a theoretical framework and practical guidelines for developing trustworthy LLMs.
This work addresses the critical issue of fluent yet audio-irrelevant hallucinations generated by automatic speech recognition (ASR) models such as Whisper, which undermine system reliability. The study presents the first systematic comparison of three reference-free hallucination detection approaches: text-based metrics, large language model prompting strategies, and probing internal states of the Whisper decoder. Furthermore, it introduces a lightweight late-fusion meta-classifier that effectively integrates signals from multiple sources. Experimental results reveal that intermediate-layer representations within the Whisper decoder alone exhibit strong discriminative power for hallucination detection. Moreover, the proposed meta-classifier, which fuses textual and internal decoder features, achieves the best trade-off between precision and recall, significantly outperforming existing methods.
研究解决了ASR系统产生无关文本的问题,通过分析两个Conformer-Large模型在不同条件下的表现,发现最终编码阶段是关键,其失败导致输出失去音频基础。
This study addresses the challenges of localizing hallucinations and retrieving supporting evidence in large language models by proposing a novel joint modeling framework for hallucination detection and input evidence alignment. Integrating encoder masked prediction, confidence estimation, and token-level alignment mechanisms, the approach enables fine-grained hallucination identification with interpretable traceability. This method overcomes limitations of existing detection techniques by precisely pinpointing hallucinated segments while associating them with high-quality input evidence. Human evaluation confirms the reliability of the alignment results, demonstrating significant improvements in the trustworthiness and transparency of generated content. Collectively, this work establishes a new paradigm for mitigating hallucinations through verifiable evidence grounding.
This work addresses the limitations of existing hallucination evaluation methods for vision-language models, which predominantly rely on model-generated samples and suffer from poor timeliness and insufficient controllability. To overcome these issues, the authors construct a dataset comprising 1,600 human-written, multilingual hallucination samples, complemented by fine-grained, span-level annotations. Through systematic comparative analysis, they demonstrate for the first time that human-authored samples significantly outperform model-generated ones in terms of distributional similarity, annotation consistency, and content controllability. These advantages enable more stable and generalizable assessment of model hallucination detection capabilities, offering a reliable alternative for benchmarking hallucinations in vision-language systems.
This work addresses the challenge that existing audio-visual large language models struggle to accurately align spoken content with visual signals, often generating semantic or temporal hallucinations. To this end, the paper introduces SVHalluc, the first benchmark specifically designed to evaluate speech–vision hallucination, systematically assessing cross-modal understanding along two dimensions: semantic consistency and temporal alignment. The benchmark employs a multi-task framework and includes comprehensive experiments covering both open-source and proprietary state-of-the-art models. Results reveal that current open-source models perform near random chance, while Gemini 2.5 Pro demonstrates significant superiority, highlighting a critical gap in existing approaches for speech–vision alignment and establishing a much-needed evaluation standard in this emerging domain.