When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the unfaithfulness problem in large language model reasoning, where models generate answers before post-hoc rationalization via shortcut learning. We propose ConfLens, a framework that tracks the dynamic evolution of confidence during generation to formally define the "premature commitment" characteristic for the first time. Furthermore, we introduce a truth-free Distributional Answer Commitment Score (DACS) to quantify belief concentration and identify shortcut reasoning, subsequently leveraging this detection signal to optimize reward model preferences. Experimental results demonstrate that our approach improves F1 scores by over 4.3% compared to baselines on mathematical and code benchmarks, significantly mitigating the misalignment between faithfulness and correctness in reward models.
πŸ“ Abstract
The reasoning trajectory of a Large Language Model (LLM) is often treated as a verbalized description of its internal reasoning. However, such trajectories can be unfaithful: a model may rely on shortcuts to reach an answer and then post-rationalize the decision with a seemingly coherent chain of thought. Detecting this shortcut reasoning is challenging because existing monitors and verifiers mainly inspect textual traces or final outcomes, rather than how the model's belief in its answer develops during generation. We introduce ConfLens, a framework that tracks the evolution of confidence in the final answer throughout reasoning. Across three shortcut reasoning settings, we observe a common pattern of premature confidence, where shortcut samples become highly confident in the final answer at early reasoning stages. Existing confidence estimation methods, however, show limited generalizability, reliability, or efficiency for detecting this behavior. We therefore propose the Distributional Answer Commitment Score (DACS), a distributional confidence estimator that measures the entropy of the model's probability distribution over answer commitment at each reasoning step. DACS captures how concentrated the model's answer belief is without requiring ground-truth answers or task-specific verifiers. We further convert ConfLens detection results into interpretable signals for reward models to reduce their preference for shortcut reasoning. Experiments on mathematical and code reasoning tasks show that ConfLens with DACS improves shortcut reasoning detection by over 4.3% F1 compared with strong baselines and reduces the mismatch between faithfulness and correctness in reward model preferences.
Problem

Research questions and friction points this paper is trying to address.

Shortcut Reasoning
Faithfulness
Premature Confidence
Large Language Models
Confidence Estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Shortcut Reasoning Detection
Confidence Evolution Tracking
Distributional Answer Commitment Score (DACS)
Premature Confidence
Reward Model Alignment
πŸ”Ž Similar Papers