Score
Designs, builds, and evaluates methods that compute and calibrate uncertainty or confidence at the individual token level produced by sequence models—e.g., token log-probabilities, conditional entropy, or posterior variance—and that aggregate token confidences to score, rank, or qualify whole responses. Also implements tokenization and decoding procedures that incorporate these confidence estimates (uncertainty-aware tokenization, confidence-based stopping or flagging) and tooling to threshold and evaluate low-confidence outputs for review.
Large language models (LLMs) exhibit reliability challenges in high-stakes domains (e.g., healthcare, law), where uncertainty arises from multiple interdependent sources—ambiguous inputs, divergent reasoning paths, parametric stochasticity, and output randomness—extending well beyond classical aleatoric/epistemic dichotomies. To address this, we propose the first four-dimensional uncertainty taxonomy for LLMs—categorizing uncertainty along input, reasoning, parameter, and prediction axes—overcoming key limitations of conventional uncertainty quantification (UQ) in dimensional coverage and computational scalability. We systematically evaluate over twenty UQ methods—including Bayesian approximation, ensemble sampling, logit calibration, attention-based analysis, and consistency verification—across real-world tasks to characterize their applicability boundaries and failure modes. Finally, we introduce a comprehensive evaluation framework balancing interpretability, robustness, and practicality, providing both theoretical foundations and actionable guidelines for deploying trustworthy LLMs in safety-critical applications.
This study addresses the “hallucination” problem in large language models (LLMs) arising from inadequate uncertainty quantification (UQ) and poor calibration. We present the first systematic survey of LLM uncertainty calibration methods and introduce the first comprehensive UQ and calibration benchmark specifically designed for LLMs. Our standardized empirical evaluation covers six major calibration approaches across two reliability-focused datasets. We propose a unified evaluation framework incorporating confidence-accuracy alignment analysis, Expected Calibration Error (ECE), and Brier Score. Results reveal that existing methods achieve only limited calibration performance; furthermore, task type, prompt engineering, and output format significantly influence uncertainty estimation quality. To foster reproducible research, we open-source our evaluation protocol and analytical toolchain. This work establishes a rigorous, community-accessible benchmark and methodological foundation for advancing LLM reliability and trustworthy AI.
This work addresses the overconfidence of large language models in mathematical question answering, which often leads to unreliable confidence estimates. The authors systematically investigate token-probability–based calibration methods, proposing a single-pass confidence estimation approach and a low-cost in-place self-verification strategy, complemented by post-hoc techniques such as Platt scaling and isotonic regression. Experimental results demonstrate that aggregating probabilities over the full output sequence effectively discriminates between correct and incorrect answers, and that multi-channel methods substantially reduce in-domain calibration error. However, calibration performance is sensitive to problem difficulty and exhibits limited transferability across different models and datasets.
This study addresses the unclear mechanisms underlying confidence estimation in large language model (LLM) reasoning and the limited effectiveness of existing probability calibration methods. To this end, it proposes Divergent Token Counting (DTC), a training-free framework that leverages Jensen-Shannon divergence to identify and count strongly divergent tokens between two models along identical decoding trajectories. The analysis reveals a significant negative correlation between divergence counts and predictive accuracy, enabling efficient uncertainty quantification in both black-box and white-box settings. Evaluated on mathematical reasoning benchmarks, DTC reduces calibration error to approximately 13%, substantially outperforming conventional baselines. Overall, this work presents a lightweight yet effective solution for assessing LLM reliability.
This study identifies a systematic misalignment between token-level output determinism in large language models (LLMs) and theoretically grounded probability distributions in probabilistic reasoning tasks: even when semantic responses are fully correct (100% accuracy), the logits over generated tokens significantly deviate from Bayesian posterior distributions. Using GPT-4.1 and DeepSeek-Chat, we analyze token-level logits across multiple prompt iterations, quantify uncertainty via entropy, and compare empirical token distributions against normative probabilistic constraints. Results demonstrate that prevailing uncertainty quantification (UQ) methods fail to ensure token-level calibration, exposing a fundamental tension between semantic correctness and probabilistic coherence. To our knowledge, this is the first empirical characterization of the “determinism–probability mismatch” phenomenon—where deterministic token selection contradicts principled uncertainty representation. Our work establishes a new benchmark for trustworthy probabilistic reasoning and motivates theoretical reexamination of UQ in autoregressive LMs.
This work addresses the susceptibility of low-bit quantized language models to generation failures caused by degeneration during inference, necessitating effective monitoring mechanisms. The authors propose a training-free decoding controller that constructs a degradation-aware alert score by integrating token-level uncertainty with explicit repetition signals. Leveraging e-process theory, they design a calibrated CUSUM sequential detector for sequence-level monitoring. The study identifies the fundamental inadequacy of centered token log-probabilities as monitoring metrics and instead adopts a historically dependent, autocorrelated, and properly calibrated alternative. Experiments on GSM8K demonstrate that the method improves the accuracy of an INT4 model from 63% to 69% (p=0.18) at a token overhead of 28%, while achieving approximately 60% precision in detecting failing trajectories and substantially mitigating verbatim degeneration.
This study addresses the challenge that local errors in large language model generation are often obscured by global confidence scores, thereby complicating calibration. To overcome this, it proposes a single-pass decoding trajectory risk localization method that pioneers modeling decoding uncertainty as trajectories. By recording token-level surprisal and predictive entropy while employing local risk operators to preserve uncertainty spikes, the approach achieves answer-level risk scoring and probability calibration without requiring labeled data. Evaluated across four tasks, the proposed method significantly outperforms nineteen baselines, reducing the Brier score to 0.137 and improving the AUROC to 0.792, which demonstrates its generalizability and effectiveness.
为解决大语言模型在决策中产生错误信息及信心错位问题,提出基于声明级别的置信度校准方法,通过分解响应并使用推理时信号进行校准。
This work addresses the challenge of interpreting the contribution of input tokens to outputs in large language model generation by proposing the first model-agnostic probabilistic attribution method. The approach models text generation as a stochastic process and leverages Bayes’ rule to infer the conditional probability of a response given a prompt. Attribution scores are defined via the logarithm of probability ratios, while conditional entropy is introduced to quantify context sensitivity and generation uncertainty. Experiments across eight mainstream models and seven prompt categories demonstrate that the method effectively identifies anomalous generations, token-sensitive regions, and unstable behaviors, substantially enhancing users’ awareness and understanding of generative uncertainty.
This study addresses the lack of uncertainty quantification and confidence calibration in the natural language outputs of activation steering, which undermines the reliability of interpretability methods. It presents the first systematic investigation of this issue, evaluating six confidence estimation techniques—including bootstrapped mode frequency and answer-token log-probabilities—across diverse prompt templates and the Qwen3 series of large language models. Using a 6,000-sample test set, the experiments demonstrate that bootstrapped mode frequency substantially outperforms conventional log-probability baselines, reducing expected calibration error (ECE) to 5.7% for Qwen3-8B and 10.3% for Qwen3.6-27B. The work further introduces a low-cost, rapid screening strategy that offers a practical solution for generating reliable activation-based explanations.
This study addresses the overlooked sensitivity of large language model (LLM) confidence calibration evaluations to measurement protocols, particularly in comparing token-level probabilities and verbalized confidence. Through systematic controlled experiments across four question-answering benchmarks, the authors investigate how protocol choices—such as answer string selection, token probability extraction methods, and conditional context formulation—affect calibration assessments for three open-source 7–8B models and their Qwen2.5 variants. The findings reveal that, under default protocols, verbalized confidence offers no significant calibration advantage over token probabilities, and that incorrect yet superficially plausible answers often receive confidence scores comparable to correct ones. The work underscores the high protocol dependence of confidence signals, advocates treating them as protocol-contingent behavioral measurements, and proposes a standardized reporting checklist to enhance reproducibility and comparability in calibration evaluation.
研究探讨了自回归语言模型中局部与全局置信度的差异及其对预测正确性和采样稳定性的影响,指出不同置信度读数不可互换。