hallucination detection and mitigation

Designs and builds detectors and mitigation pipelines that identify and reduce hallucinated model outputs across modalities and output granularities — including training-free, unsupervised, label-free and white-box detectors that use hidden-state or trace-based trust scoring, span-level and visual-region detection, and citation-focused methods that compare, normalize, and verify references (including legal citation formats). Implements filters and suppression strategies, integrates detection into extraction or rollout pipelines, measures hallucination rates and distinct failure modes, evaluates mitigation effectiveness across models, and analyzes rollouts to flag low-coverage state-actions and estimate failure likelihood.

hallucinationdetectionandmitigation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$206K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Large language models (LLMs) suffer from pervasive hallucination—generating syntactically fluent yet factually incorrect or unsupported content—undermining their reliability and practical utility. This paper presents a systematic survey of hallucination causes, detection techniques, and mitigation strategies. We propose the first comprehensive taxonomy spanning the entire LLM lifecycle—data curation, model architecture, and inference—alongside a dual-dimensional classification framework (distinguishing detection vs. mitigation across methodological layers: token-, sequence-, and system-level). Our analysis exposes fundamental limitations in existing approaches regarding benchmark compatibility, metric validity, and cross-domain generalizability. Furthermore, we conduct a multi-faceted evaluation of mainstream hallucination benchmarks to assess their theoretical soundness and empirical rigor. The study establishes a principled, verifiable framework for developing trustworthy LLMs, offering concrete methodological guidance toward reliable and accountable AI systems.

Creating strategies to mitigate false information from language modelsDeveloping methods to detect hallucinated content in generated textInvestigating causes of factual inaccuracies in large language models

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the susceptibility of large language models to hallucinations in high-stakes domains such as finance and law, which undermines output reliability. The authors propose a root-cause-aware continuous improvement framework that categorizes hallucination sources into three types: model-induced, data-related, and context-driven. By integrating techniques including uncertainty estimation, reasoning consistency analysis, knowledge anchoring, and confidence calibration, the framework establishes a closed-loop mechanism for hierarchical detection and targeted mitigation. This approach shifts the paradigm from generic post-hoc fixes to precise, cause-specific governance. Evaluated on financial data extraction tasks, the method significantly enhances both generation accuracy and trustworthiness, offering a scalable solution for deploying reliable AI systems in regulation-sensitive scenarios.

FactualityGenerative AIHallucination

This work addresses the challenge of hallucinations in large language models (LLMs), which hinder their reliable deployment in high-stakes applications. The authors propose a novel hallucination detection method that leverages gradient patterns across model layers during a single forward–backward pass. They discover, for the first time, that over 97% of discriminative gradient signals are concentrated in the final five layers of the model. Building upon this insight, they develop an efficient and interpretable unified framework capable of simultaneously identifying hallucinated content and predicting when the model should abstain from answering. Experimental results across multiple question-answering benchmarks demonstrate that the proposed approach significantly outperforms baseline methods based on confidence scores or sampling strategies, achieving high performance with minimal computational overhead.

gradient analysishallucination detectionLarge Language Models

This work addresses the critical challenge of hallucination-induced cascading failures in GUI agents during real-world deployment, where existing approaches lack fine-grained diagnosis, reliable evaluation, and efficient mitigation mechanisms. To this end, the paper proposes the first hallucination governance framework tailored for GUI environments. It introduces an empirically grounded hallucination taxonomy, a three-stage calibration and evaluation pipeline, and a lightweight closed-loop structured reasoning module augmented with a cold-start post-training strategy. Remarkably, the approach achieves a significant reduction in hallucination rates using only 9K training samples, substantially enhancing the agent’s environmental grounding and operational fidelity without requiring extensive computational resources.

diagnosisevaluationGUI agents

In enterprise settings, large language models (LLMs) suffer from hallucination due to limited context windows and outdated knowledge; existing mitigation strategies—such as gold-standard QA repositories or secondary verification models—are costly and lack formal guarantees of correctness. This paper proposes an interactive, visualization-enabled knowledge graph framework for hallucination detection: LLM-generated assertions are dynamically linked to proprietary knowledge sources to construct a structured truth-graph, supporting confidence scoring, provenance tracing, and human-in-the-loop feedback. Our key contributions lie in the integration of adaptive knowledge graph construction, interpretable natural language understanding (NLU), and human–AI collaborative diagnosis—enabling real-time identification and auditable verification of hallucinated content. Experiments demonstrate significant improvements in LLM response trustworthiness and reliability under constrained context and knowledge inconsistency, while establishing a sustainable, feedback-driven optimization loop.

Detects hallucinations in LLMs using visual knowledge graphsEnables human feedback to improve model reliability continuouslyLinks model assertions to truth sources for user verification

Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback

Apr 22, 2024
WX
Wenyi Xiao
🏛️ Zhejiang University | Alibaba Group

To address hallucination in large vision-language models (LVLMs) during image captioning—caused by image-text misalignment—this paper proposes a fine-grained, AI-feedback-driven detection and mitigation framework. Methodologically, it introduces the first sentence-level, multi-type hallucination detector identifying object-, attribute-, and relation-level inconsistencies; designs hallucination-severity-aware direct preference optimization (HSA-DPO) to close a detection–rewriting–preference-learning loop; and operates entirely without human annotations or reliance on black-box foundation models. The key contribution is a lightweight, transferable end-to-end solution that significantly improves LVLM reliability across multiple benchmarks: hallucination detection F1 score increases by 12.6%, and image-text alignment of generated captions improves by 23.4%. This work establishes a novel paradigm for enhancing LVLM robustness and factual consistency.

Large Visual Language ModelsMisalignment ErrorsReliability in Practical Applications

Latest Papers

What's happening recently
View more

This study addresses the potential overestimation of hallucination detection performance due to dataset construction artifacts—particularly prompt leakage—in existing benchmarks. Through a systematic evaluation of 22 detection methods, 12 open-source models, and 6 corpora, the work quantifies, for the first time, the extent to which such artifacts inflate reported results. To enable reliable real-time hallucination detection, the authors propose DRIFT, a supervised probing method based on transitions in upper-layer hidden states. Experimental findings reveal that, once prompt leakage is controlled, most existing approaches perform near random chance, with only SAPLMA and DRIFT demonstrating consistent effectiveness across diverse settings. These results indicate that current progress in hallucination detection has been substantially overstated and establish a more trustworthy evaluation framework for real-world applications.

benchmark artifactsevaluation biashallucination detection

This work addresses the challenge of detecting hallucinated citations generated by large language models in academic writing, a problem exacerbated by existing methods that rely on fragile parsing or incomplete retrieval and thus lack fine-grained discriminative capability. The authors propose CiteTracer, the first multi-agent cascaded detection framework capable of field-level citation verification, reframing hallucination detection as a 12-class truthfulness classification task. Integrating structured parsing (from PDFs and BibTeX), multi-source evidence retrieval (via academic search engines, web search, and URL scraping), and an expert routing mechanism, CiteTracer achieves 97.1% overall accuracy on a benchmark comprising 2,450 synthetic and 957 real-world hallucinated citations, with F1 scores of 97.0, 95.8, and 98.5 across three critical categories—significantly advancing the detection of ambiguous and fabricated references.

bibliographic verificationcitation hallucinationfake references

Enterprise AI systems struggle to gain trust due to hallucinations—confident yet incorrect outputs from large language models—a risk that existing approaches fail to eliminate. This work proposes HALO, a novel architecture that treats hallucination as a manageable system failure mode and enforces a “zero-hallucination” guarantee through six coordinated defense layers: retrieval-constrained generation, deterministic execution constraints, multi-signal validation (integrating LLM-based discriminators with source document alignment), calibrated refusal, end-to-end traceability, and a continuous monitoring feedback loop. Evaluated on insurance claim information extraction, HALO delivers high-fidelity outputs, effectively blocks hallucinations, and provides early warnings of system drift, substantially enhancing the trustworthiness of enterprise AI deployments.

enterprise AIfactualitygrounded generation

This work addresses the critical issue of large language models (LLMs) hallucinating non-existent software packages during code generation, which introduces significant risks of typosquatting-based supply chain attacks. Existing mitigation strategies either incur high computational costs or compromise the model’s general capabilities. To overcome these limitations, the authors propose an Adaptive Unlearning (AU) framework—the first post-training intervention that operates without human annotations or predefined forget sets. AU employs an adaptive discovery mechanism to continuously identify emerging hallucination scenarios and integrates a hybrid token-level optimization objective to precisely suppress package-related hallucinations using only model-generated data, while preserving valid outputs. The method reduces package hallucination rates by 81%, substantially narrowing the attack surface, and maintains competitive performance on standard code generation benchmarks.

code generationhallucinationlarge language models

This work addresses the prevalent issue of object hallucination in large vision-language models, where generated image captions often contain factually incorrect content inconsistent with the input image, thereby undermining model reliability. The authors propose a training-free, test-time hallucination mitigation method (TTH) that leverages a zero-shot multimodal classifier to produce image-grounded auxiliary logits for candidate object tokens. These logits are fused at the token level with the model’s original outputs, guided by an entropy-based weighting strategy to refine predictions. Notably, this approach achieves effective and general-purpose hallucination control without requiring multiple decoding passes or architectural modifications to the underlying model. Extensive experiments demonstrate significant improvements in accuracy and robustness across diverse vision-language models and benchmarks, while fully preserving the models’ pretrained knowledge.

Large Vision-Language ModelsNon-factual GenerationObject Hallucination

Hot Scholars

HY

Hao Yin

Meta Platforms Inc.
Wireless communicationOptimizationMachine Learning
MG

Mor Geva

Tel Aviv University, Google Research
Natural Language Processing
XH

Xuming Hu

Assistant Professor, HKUST(GZ) / HKUST
Natural Language ProcessingLarge Language Model
IR

Imran Razzak

MBZUAI, Abu Dhabi
Human-Centered AIMedical Image AnalysisMedical Artificial IntelligenceComputational Biology
DY

Dawei Yin

Senior Director, Head of Search Science at Baidu
Machine LearningWeb MiningData Mining