per-token uncertainty estimation

Designs, builds, and evaluates methods that compute and calibrate uncertainty or confidence at the individual token level produced by sequence models—e.g., token log-probabilities, conditional entropy, or posterior variance—and that aggregate token confidences to score, rank, or qualify whole responses. Also implements tokenization and decoding procedures that incorporate these confidence estimates (uncertainty-aware tokenization, confidence-based stopping or flagging) and tooling to threshold and evaluate low-confidence outputs for review.

per-tokenuncertaintyestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.41
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Comparing Uncertainty Measurement and Mitigation Methods for Large Language Models: A Systematic Review

Apr 25, 2025
TA
Toghrul Abbasli
🏛️ Tsinghua University | Institute of High Performance Computing | Agency for Science, Technology and Research | King Abdullah University of Science and Technology | Zhongguancun Laboratory

This study addresses the “hallucination” problem in large language models (LLMs) arising from inadequate uncertainty quantification (UQ) and poor calibration. We present the first systematic survey of LLM uncertainty calibration methods and introduce the first comprehensive UQ and calibration benchmark specifically designed for LLMs. Our standardized empirical evaluation covers six major calibration approaches across two reliability-focused datasets. We propose a unified evaluation framework incorporating confidence-accuracy alignment analysis, Expected Calibration Error (ECE), and Brier Score. Results reveal that existing methods achieve only limited calibration performance; furthermore, task type, prompt engineering, and output format significantly influence uncertainty estimation quality. To foster reproducible research, we open-source our evaluation protocol and analytical toolchain. This work establishes a rigorous, community-accessible benchmark and methodological foundation for advancing LLM reliability and trustworthy AI.

Assessing and quantifying uncertainty in Large Language ModelsEvaluating effectiveness of uncertainty measurement methods for LLMsProviding a benchmark for comparing calibration techniques in LLMs

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the overconfidence of large language models in mathematical question answering, which often leads to unreliable confidence estimates. The authors systematically investigate token-probability–based calibration methods, proposing a single-pass confidence estimation approach and a low-cost in-place self-verification strategy, complemented by post-hoc techniques such as Platt scaling and isotonic regression. Experimental results demonstrate that aggregating probabilities over the full output sequence effectively discriminates between correct and incorrect answers, and that multi-channel methods substantially reduce in-domain calibration error. However, calibration performance is sensitive to problem difficulty and exhibits limited transferability across different models and datasets.

calibrationconfidence estimationlarge language models

This study addresses the unclear mechanisms underlying confidence estimation in large language model (LLM) reasoning and the limited effectiveness of existing probability calibration methods. To this end, it proposes Divergent Token Counting (DTC), a training-free framework that leverages Jensen-Shannon divergence to identify and count strongly divergent tokens between two models along identical decoding trajectories. The analysis reveals a significant negative correlation between divergence counts and predictive accuracy, enabling efficient uncertainty quantification in both black-box and white-box settings. Evaluated on mathematical reasoning benchmarks, DTC reduces calibration error to approximately 13%, substantially outperforming conventional baselines. Overall, this work presents a lightweight yet effective solution for assessing LLM reliability.

CalibrationChain-of-ThoughtLarge Language Models

Certain but not Probable? Differentiating Certainty from Probability in LLM Token Outputs for Probabilistic Scenarios

Nov 01, 2025
AT
Autumn Toney-Wails
🏛️ SciTech Strategies, Inc. | Georgetown University

This study identifies a systematic misalignment between token-level output determinism in large language models (LLMs) and theoretically grounded probability distributions in probabilistic reasoning tasks: even when semantic responses are fully correct (100% accuracy), the logits over generated tokens significantly deviate from Bayesian posterior distributions. Using GPT-4.1 and DeepSeek-Chat, we analyze token-level logits across multiple prompt iterations, quantify uncertainty via entropy, and compare empirical token distributions against normative probabilistic constraints. Results demonstrate that prevailing uncertainty quantification (UQ) methods fail to ensure token-level calibration, exposing a fundamental tension between semantic correctness and probabilistic coherence. To our knowledge, this is the first empirical characterization of the “determinism–probability mismatch” phenomenon—where deterministic token selection contradicts principled uncertainty representation. Our work establishes a new benchmark for trustworthy probabilistic reasoning and motivates theoretical reexamination of UQ in autoregressive LMs.

Assessing token-level probability divergence from expected theoretical distributionsDifferentiating certainty from probability in LLM token outputsEvaluating alignment with theoretical probability distributions in probabilistic scenarios

This work addresses the susceptibility of low-bit quantized language models to generation failures caused by degeneration during inference, necessitating effective monitoring mechanisms. The authors propose a training-free decoding controller that constructs a degradation-aware alert score by integrating token-level uncertainty with explicit repetition signals. Leveraging e-process theory, they design a calibrated CUSUM sequential detector for sequence-level monitoring. The study identifies the fundamental inadequacy of centered token log-probabilities as monitoring metrics and instead adopts a historically dependent, autocorrelated, and properly calibrated alternative. Experiments on GSM8K demonstrate that the method improves the accuracy of an INT4 model from 63% to 69% (p=0.18) at a token overhead of 28%, while achieving approximately 60% precision in detecting failing trajectories and substantially mitigating verbatim degeneration.

chain-of-thought degradationdecoder monitoringgeneration reliability

This study addresses the challenge that local errors in large language model generation are often obscured by global confidence scores, thereby complicating calibration. To overcome this, it proposes a single-pass decoding trajectory risk localization method that pioneers modeling decoding uncertainty as trajectories. By recording token-level surprisal and predictive entropy while employing local risk operators to preserve uncertainty spikes, the approach achieves answer-level risk scoring and probability calibration without requiring labeled data. Evaluated across four tasks, the proposed method significantly outperforms nineteen baselines, reducing the Brier score to 0.137 and improving the AUROC to 0.792, which demonstrates its generalizability and effectiveness.

confidence estimationdecoding uncertaintygeneration calibration

Latest Papers

What's happening recently
View more

This work addresses the challenge of interpreting the contribution of input tokens to outputs in large language model generation by proposing the first model-agnostic probabilistic attribution method. The approach models text generation as a stochastic process and leverages Bayes’ rule to infer the conditional probability of a response given a prompt. Attribution scores are defined via the logarithm of probability ratios, while conditional entropy is introduced to quantify context sensitivity and generation uncertainty. Experiments across eight mainstream models and seven prompt categories demonstrate that the method effectively identifies anomalous generations, token-sensitive regions, and unstable behaviors, substantially enhancing users’ awareness and understanding of generative uncertainty.

interpretabilitylarge language modelsprobabilistic attribution

This study addresses the lack of uncertainty quantification and confidence calibration in the natural language outputs of activation steering, which undermines the reliability of interpretability methods. It presents the first systematic investigation of this issue, evaluating six confidence estimation techniques—including bootstrapped mode frequency and answer-token log-probabilities—across diverse prompt templates and the Qwen3 series of large language models. Using a 6,000-sample test set, the experiments demonstrate that bootstrapped mode frequency substantially outperforms conventional log-probability baselines, reducing expected calibration error (ECE) to 5.7% for Qwen3-8B and 10.3% for Qwen3.6-27B. The work further introduces a low-cost, rapid screening strategy that offers a practical solution for generating reliable activation-based explanations.

activation oraclesconfidence calibrationinterpretability

This study addresses the overlooked sensitivity of large language model (LLM) confidence calibration evaluations to measurement protocols, particularly in comparing token-level probabilities and verbalized confidence. Through systematic controlled experiments across four question-answering benchmarks, the authors investigate how protocol choices—such as answer string selection, token probability extraction methods, and conditional context formulation—affect calibration assessments for three open-source 7–8B models and their Qwen2.5 variants. The findings reveal that, under default protocols, verbalized confidence offers no significant calibration advantage over token probabilities, and that incorrect yet superficially plausible answers often receive confidence scores comparable to correct ones. The work underscores the high protocol dependence of confidence signals, advocates treating them as protocol-contingent behavioral measurements, and proposes a standardized reporting checklist to enhance reproducibility and comparability in calibration evaluation.

confidence calibrationlarge language modelsprotocol sensitivity

Hot Scholars

TF

Tianyu Fu

Ph.D at Tsinghua University
efficient AILLMsparse computation
XJ

Xiaojun Jia

Nanyang Technological University
Explainable AIRobust AIEfficient AI
SP

Shirui Pan

Professor, ARC Future Fellow, FQA, Director of TrustAGI Lab, Griffith University
Data MiningMachine LearningGraph Neural NetworksTrustworthy AI
WL

Weiming Lu

Zhejiang University
Natural Language ProcessingLarge Language ModelsAGI