Score
Designs and executes analyses and metrics to detect, quantify, and characterize hallucinations in model outputs, including categorizing severity, discovering recurring pattern types, measuring cross-model vocabulary and instance overlap, and identifying conditions under which hallucinations occur; builds visualizations, summary statistics, and comparative tests to support diagnosis and root-cause investigation.
Hallucinations in large language and vision models critically undermine the reliability and safety of generative AI deployments, yet existing research lacks a unified, systematic understanding of their root causes. This project introduces the first cross-modal, multi-level hallucination analysis framework, jointly considering task- and modality-specific dimensions. It identifies hallucinations as arising from the synergistic interaction between data distribution shifts and inherited model biases, and characterizes their propagation mechanisms across the full model lifecycle—training, inference, and deployment. Leveraging hierarchical classification, cross-modal comparative analysis, and large-scale mechanistic attribution studies, we establish a unified theoretical model covering both textual and visual modalities. Our framework provides a generalizable foundation for hallucination attribution, enabling the design of robust, interpretable mitigation strategies and significantly enhancing the trustworthiness and generalization capability of generative AI systems.
Hallucination—generating factually incorrect or fabricated content that appears plausible—severely undermines the reliability and trustworthiness of large language models (LLMs). Method: This work establishes the first universal theoretical framework for LLM hallucination, formally defining its essence and proving its intrinsic inevitability within computable models. It innovatively distinguishes *intrinsic* from *extrinsic* hallucination and rigorously clarifies the conceptual boundaries between *factual accuracy* and *faithfulness*. A fine-grained taxonomy is developed, covering cross-modal and multi-task scenarios. The analysis integrates theoretical modeling, classification-based formalization, and empirical validation—including data provenance tracing, logical consistency checking, benchmark evaluation, and human-subject experiments. Contribution/Results: The study systematically uncovers root causes and human perception mechanisms of hallucination. It releases an open-source evaluation benchmark and an online resource platform, providing a unified theoretical foundation and reusable toolset for hallucination detection, mitigation, and governance.
Large language models (LLMs) are prone to hallucinations—such as omitting critical information or fabricating details—when generating structured summaries of software bug reports, potentially misleading developers and undermining tool reliability. This work presents the first systematic analysis of such hallucinations from a section-aware perspective. The authors introduce a controllable synthetic hallucination injection mechanism to construct a benchmark dataset and propose a multi-task joint detection framework that simultaneously predicts whether a report contains hallucinations, localizes the affected sections, and identifies the hallucination type. Experiments on the BugsRepo dataset demonstrate that the best-performing model achieves Macro-F1 scores of 0.89, 0.83, and 0.84 at the report, section, and hallucination-type levels, respectively, effectively uncovering prevalent hallucination patterns and underlying model failure mechanisms.
Existing geometric hallucination detection metrics struggle to distinguish specific hallucination types in the absence of ground truth and are highly sensitive to domain shifts. This work addresses these limitations by constructing a synthetic dataset to systematically evaluate the capacity of various geometric statistics to capture key hallucination attributes—such as output correctness, relevance, and coherence—and reveals that different metrics align with distinct hallucination types. Furthermore, the study proposes a simple yet effective normalization strategy that substantially mitigates the impact of domain shift. Experimental results demonstrate that, under multi-domain settings, the proposed approach improves AUROC by 34 percentage points, significantly enhancing the cross-domain robustness of geometric hallucination detection metrics.
This study is the first to systematically reveal the prevalence of hallucinations in large language models (LLMs) for natural language generation from code changes—specifically, commit message and code review comment generation—finding factual inaccuracies in approximately 50% of review comments and 20% of commit messages. Method: To address this, we propose the first multi-metric hallucination detection framework tailored to code-change scenarios, which jointly models model confidence, gradient-based feature attribution, and semantic similarity—enabling efficient, fine-tuning-free detection at inference time. Contribution/Results: Extensive experiments demonstrate that our method significantly outperforms single-metric baselines across diverse LLMs and code-change datasets. It provides a practical, empirically validated technical pathway to enhance the factual reliability and trustworthiness of code-related NLG systems, establishing foundational evidence for hallucination mitigation in software engineering AI applications.
Foundation models frequently exhibit “hallucinations” during autonomous decision-making, leading to high-risk misjudgments—yet no formal definition or systematic characterization of hallucination exists for decision tasks. Method: This work introduces the first task-specific definition of hallucination in decision-making and establishes a cross-task, scalable hallucination taxonomy; proposes a synergistic framework integrating uncertainty quantification with hallucination detection; and develops a joint detection–decision evaluation paradigm grounded in systematic survey analysis, probabilistic uncertainty modeling, and decision-chain interpretability. Results: We present the first comprehensive landscape of hallucination detection techniques tailored to decision contexts, contributing seven actionable implementation guidelines and identifying five critical research gaps—thereby advancing the safe deployment of trustworthy foundation models in high-stakes domains such as healthcare and transportation.
This work addresses the critical challenge of hallucination in large language models during tool-augmented reasoning, which often leads to incorrect tool selection, erroneous parameter specification, or unjustified tool bypassing—thereby compromising system reliability. The authors propose a novel real-time hallucination detection method that leverages internal model representations from a single forward pass, eliminating the need for additional inference steps or external validation. By deploying a lightweight classifier to analyze activation patterns within the same inference cycle, the approach efficiently identifies hallucinations at both the tool-selection and parameter-specification levels. Evaluated across diverse domains, the method achieves up to 86.4% detection accuracy, substantially outperforming existing techniques while introducing negligible inference latency, thus significantly enhancing both the safety and efficiency of deployed systems.
In enterprise settings, large language models (LLMs) suffer from hallucination due to limited context windows and outdated knowledge; existing mitigation strategies—such as gold-standard QA repositories or secondary verification models—are costly and lack formal guarantees of correctness. This paper proposes an interactive, visualization-enabled knowledge graph framework for hallucination detection: LLM-generated assertions are dynamically linked to proprietary knowledge sources to construct a structured truth-graph, supporting confidence scoring, provenance tracing, and human-in-the-loop feedback. Our key contributions lie in the integration of adaptive knowledge graph construction, interpretable natural language understanding (NLU), and human–AI collaborative diagnosis—enabling real-time identification and auditable verification of hallucinated content. Experiments demonstrate significant improvements in LLM response trustworthiness and reliability under constrained context and knowledge inconsistency, while establishing a sustainable, feedback-driven optimization loop.
Existing hallucination detection methods focus narrowly on factual consistency, overlooking potential creative value and struggling to balance accuracy with creativity across diverse scientific tasks. Method: We propose HIC-Bench—a novel benchmark framework that systematically distinguishes *Intelligent Hallucination* (IH) from *Defective Hallucination* (DH). It evaluates both dimensions across ten open-ended scientific innovation tasks using a dual-axis metric: creativity (integrating Torrance Tests of Creative Thinking with hallucination-specific dimensions) and factual deviation. Innovations include an IH/DH binary classification paradigm, Dynamic Hallucination Prompting (DHP), a multidimensional metrics matrix, cross-disciplinary task design, ensemble evaluation by multiple LLMs, and human validation. Results: IH and DH exhibit a nonlinear relationship; creativity and factual accuracy can be jointly optimized; and hallucinations—when appropriately structured—can serve as catalysts for scientific innovation.
This work addresses the limitation of existing hallucination evaluation methods, which predominantly focus on final outputs while overlooking hallucinations in intermediate reasoning trajectories within multi-agent industrial workflows. To bridge this gap, the authors introduce the Trajel dataset and an accompanying evaluation framework, leveraging expert-annotated trajectories from AssetOpsBench to establish a five-dimensional taxonomy of hallucinations—encompassing factuality, referentiality, logicality, procedural coherence, and scope adherence—enabling, for the first time, fine-grained identification of concurrent, multi-type trajectory-level hallucinations. By deploying supervised detection models across subtask, trajectory, and long-context levels, the study reveals that nearly half of hallucinated trajectories exhibit multiple hallucination types, a complexity largely missed by current benchmarks; notably, trajectory-aware detection substantially outperforms conventional post-hoc verification approaches.
This study addresses the critical issue of hallucinations in medical imaging AI—such as fabricated anatomical structures, left-right confusion, or erroneous measurements—which can lead to misdiagnosis and inappropriate treatment. The work proposes the first cross-modal hallucination taxonomy encompassing the entire imaging pipeline, systematically evaluating hallucination tendencies in both general-purpose and medical-specific foundation models. It integrates physical constraints, chain-of-thought prompting, and human-in-the-loop mechanisms to develop a detection and mitigation strategy aligned with FDA’s full lifecycle regulatory requirements. Key findings reveal that medical-specific models, due to overfitting, are paradoxically more prone to hallucinations; that combining multiple mitigation strategies effectively addresses diverse failure modes; and that radiologist review remains essential for safe clinical deployment.