Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs
This study addresses the unclear mechanisms underlying confidence estimation in large language model (LLM) reasoning and the limited effectiveness of existing probability calibration methods. To this end, it proposes Divergent Token Counting (DTC), a training-free framework that leverages Jensen-Shannon divergence to identify and count strongly divergent tokens between two models along identical decoding trajectories. The analysis reveals a significant negative correlation between divergence counts and predictive accuracy, enabling efficient uncertainty quantification in both black-box and white-box settings. Evaluated on mathematical reasoning benchmarks, DTC reduces calibration error to approximately 13%, substantially outperforming conventional baselines. Overall, this work presents a lightweight yet effective solution for assessing LLM reliability.