🤖 AI Summary
This study addresses the severe verbal overconfidence problem in reasoning language models, which is difficult to calibrate and ineffective for guiding answer selection. By distinguishing between "certainty" and "correctness," this work proposes a probe-guided self-distillation method based on internal model states. Specifically, it employs linear probes on hidden layers to extract internal confidence signals, using supervised fine-tuning to drive the model toward producing well-calibrated verbalized confidence without requiring reinforcement learning, yet achieving RL-level calibration performance. Experiments demonstrate that Qwen3-14B reduces its in-domain expected calibration error (ECE) to 0.024. Furthermore, the proposed approach exhibits significantly superior out-of-distribution generalization compared to post-hoc recalibration baselines and effectively supports downstream applications such as weighted voting.
📝 Abstract
Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer, not how likely the answer is to be right. Post-hoc rescaling therefore fits one distribution but rarely transfers. We look inside the model instead. On factual question answering, a linear probe on the hidden state between the chain of thought and the answer is substantially better calibrated: its expected calibration error is 5 to 38 times lower than that of the verbalized score across four benchmarks and two model families. However, when used to pick among N sampled answers, that same probe nearly ties majority voting yet falls far short of the oracle. Internal states answer"how certain am I"well and"which answer is right"poorly, so the signal should be reported as a confidence rather than used to select answers. As a result, we introduce probe-guided self-distillation (Probe-SD): score a model's own sampled traces with the probe, overwrite the confidence each trace states, and finetune the base checkpoint of the same family, so nothing but the model itself remains at test time. On Qwen3-14B, Probe-SD cuts ECE from 0.178 to 0.024 in-domain and from 0.542 to 0.113 out-of-domain, where it also beats post-hoc recalibration and self-consistency distillation. The resulting confidence is well-calibrated and useful for weighted voting, behaviors previously attributed to online RL, here obtained with supervised finetuning alone.