The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the hallucination problem in language models, where they confidently generate responses despite lacking knowledge, a phenomenon whose internal decision mechanisms remain unexplored. We employ causal gating techniques to localize sparse circuits responsible for commit-abstain decisions, revealing an internal dynamic pattern characterized by early commitment accumulation and insufficient late-stage correction. Building on this insight, we introduce a lightweight policy network that processes circuit activation signals to optimize decision-making. Experimental results demonstrate that our approach improves decision accuracy by 12.2 points and reduces erroneous abstention rates by 2.5 times. Furthermore, the method exhibits strong generalization across multiple benchmarks and large-scale models ranging from 27B to 35B parameters.
📝 Abstract
Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to answer reliably. A large body of existing work mitigates hallucination through detection or abstention mechanisms, but leaves open how models internally arrive at the decision to commit or abstain in the first place. We study this decision through mechanistic analysis, framing hallucination as unsupported commitment: the model commits despite exhibiting signals of unanswerability. Using causal gating, we identify a Commit-Abstain Circuit (CAC), a sparse, causally localised subset of attention heads and MLP sublayers underlying this decision. Across ten LMs (3B-14B) from five families and three benchmarks, the CAC exhibits a recurring accumulate-yet-undercorrect pattern: commitment-promoting components build up commitment in earlier layers, while abstention-promoting components act later as corrective signals that are often insufficient to overturn the accumulated commitment. Building on this finding, a lightweight policy trained on CAC activations improves decision accuracy by 12.2 points over the model's intrinsic commit-abstain margin, reduces false abstentions by 2.5 times, transfers to unseen benchmarks, and extends to larger models (27B-35B). The CAC is both diagnostic, clarifying how models overcommit, and practical, enabling improved abstention decisions.
Problem

Research questions and friction points this paper is trying to address.

Hallucination
Abstention
Language Models
Mechanistic Interpretability
Commit-Abstain Decision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mechanistic Interpretability
Commit-Abstain Circuit
Hallucination Mitigation
Causal Gating
Abstention