PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear mechanisms of hallucination propagation during multi-stage reasoning in large language models (LLMs) and the absence of response-level correction evaluation. We construct a cross-domain benchmark and introduce a novel structured evaluation paradigm for response-level behaviors, independent of final answer correctness. Through controlled experiments and trajectory behavior classification, we quantify post-hallucination reasoning trajectories using metrics such as compliance and avoidance, and train lightweight predictive models. Our findings reveal that successful recovery is rare and highly dependent on belief updating, while the predictor achieves an AUROC of 0.847. This work provides a new evaluation framework and empirical evidence for hallucination correction in LLMs.
📝 Abstract
Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning. Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood. In this work, we introduce PHRBench, a controlled benchmark for behaviorally structured PHR across four domains and 18 large language models. PHRBench characterizes each reasoning trajectory independently of final-answer correctness through Hallucination Compliance, Hallucination Avoidance, and Heuristic Correction, and defines an insightful trajectory as successful correction that ultimately reaches the correct answer. Across 4820 controlled instances, we find that successful recovery remains relatively rare and is associated with more frequent belief updates along the reasoning trajectory. We further find that properties of the hallucinated prompt contain substantial predictive signal for successful recovery, with a lightweight predictor achieving an AUROC of 0.847. These findings provide a behavioral view of post-hallucination reasoning, characterizing how LLMs resolve erroneous context and when successful recovery is likely to occur.
Problem

Research questions and friction points this paper is trying to address.

Post-Hallucination Reasoning
Hallucination
Large Language Models
Behavioral Evaluation
Reasoning Trajectory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-Hallucination Reasoning
PHRBench
Behavioral Evaluation
Reasoning Trajectory
Hallucination Recovery
🔎 Similar Papers
No similar papers found.
L
Linghao Meng
National University of Singapore
F
Feng He
Independent Researcher
X
Xuan Yang
National University of Singapore
Junyuan Mao
Junyuan Mao
National University of Singapore
LLMAI AgentAI for Healthcare
P
Pinze Ren
Tsinghua University
D
Deqing Mu
Johns Hopkins University
H
Hesen Yang
National University of Singapore
Qiankun Li
Qiankun Li
Research Fellow@NTU, Ph.D.@USTC
MLLMAI4HealthComputer VisionPattern RecognitionTrustworthy AI