Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear mechanisms underlying confidence estimation in large language model (LLM) reasoning and the limited effectiveness of existing probability calibration methods. To this end, it proposes Divergent Token Counting (DTC), a training-free framework that leverages Jensen-Shannon divergence to identify and count strongly divergent tokens between two models along identical decoding trajectories. The analysis reveals a significant negative correlation between divergence counts and predictive accuracy, enabling efficient uncertainty quantification in both black-box and white-box settings. Evaluated on mathematical reasoning benchmarks, DTC reduces calibration error to approximately 13%, substantially outperforming conventional baselines. Overall, this work presents a lightweight yet effective solution for assessing LLM reliability.
📝 Abstract
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear. Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence. We introduce Divergent Token Confidence (DTC), a framework that estimates confidence by counting tokens at which two models strongly disagree during decoding. DTC identifies these divergent tokens using the Jensen-Shannon divergence between next-token distributions evaluated along the same reasoning trajectory. We find that their count is almost negatively associated with answer accuracy, thereby serving as a simple yet effective signal for uncertainty quantification. DTC supports both white-box and black-box evaluation using auxiliary models, without explicit training and affecting the generation process. Experiments across multiple model families and six mathematical benchmarks demonstrate improved calibration over probability-based and verbalized baselines. Under white-box evaluation, the count-only estimator achieves an average expected calibration error of 13.0%, compared with 32.7%-42.4% for standard full-sequence confidence methods. In black-box settings, it also improves calibration over the original verbalized scores. For example, mean expected calibration error falls from 32.1%-40.2% to 13.7%-16.3% on DeepSeek-V3.2. These findings provide new insights for improving reasoning uncertainty quantification in large language models. The code is released at https://github.com/szu-tera/DTC.git.
Problem

Research questions and friction points this paper is trying to address.

Uncertainty Quantification
Reasoning Confidence
Large Language Models
Calibration
Chain-of-Thought
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uncertainty Quantification
Divergent Token Confidence
Jensen-Shannon Divergence
Chain-of-Thought Reasoning
Calibration
F
Feiyang Li
College of Computer Science and Software Engineering, Shenzhen University; RayNeo.AI
S
Shengjing Liu
College of Computer Science and Software Engineering, Shenzhen University
Q
Qi Zhan
College of Computer Science and Software Engineering, Shenzhen University
S
Sijie Cheng
RayNeo.AI; Tsinghua University
W
Weiqing Wang
College of Computer Science and Software Engineering, Shenzhen University
H
Hongwen Chen
College of Computer Science and Software Engineering, Shenzhen University
Y
Yuxuan Yang
College of Computer Science and Software Engineering, Shenzhen University
Wen Wang
Wen Wang
RayNeo.AI; Behavioral and Spatial AI Lab, Peking University & Tongji University
Yile Wang
Yile Wang
Shenzhen University
Natural Language Processing
Hui Huang
Hui Huang
Chair Professor and CS Dean, Shenzhen University
GraphicsGeometryPointsShapesImages