Score
Designs, implements, and evaluates methods and protocols that convert raw outputs or scores from large language models into well-calibrated probability estimates and decision rules, including calibration procedures and operational protocols. Analyzes and proves finite-sample concentration bounds and asymptotic properties (e.g., asymptotic Bayes risk) of those calibrated estimators to ensure reliable uncertainty quantification and decision-making.
Large language models (LLMs) exhibit reliability challenges in high-stakes domains (e.g., healthcare, law), where uncertainty arises from multiple interdependent sources—ambiguous inputs, divergent reasoning paths, parametric stochasticity, and output randomness—extending well beyond classical aleatoric/epistemic dichotomies. To address this, we propose the first four-dimensional uncertainty taxonomy for LLMs—categorizing uncertainty along input, reasoning, parameter, and prediction axes—overcoming key limitations of conventional uncertainty quantification (UQ) in dimensional coverage and computational scalability. We systematically evaluate over twenty UQ methods—including Bayesian approximation, ensemble sampling, logit calibration, attention-based analysis, and consistency verification—across real-world tasks to characterize their applicability boundaries and failure modes. Finally, we introduce a comprehensive evaluation framework balancing interpretability, robustness, and practicality, providing both theoretical foundations and actionable guidelines for deploying trustworthy LLMs in safety-critical applications.
This study addresses the “hallucination” problem in large language models (LLMs) arising from inadequate uncertainty quantification (UQ) and poor calibration. We present the first systematic survey of LLM uncertainty calibration methods and introduce the first comprehensive UQ and calibration benchmark specifically designed for LLMs. Our standardized empirical evaluation covers six major calibration approaches across two reliability-focused datasets. We propose a unified evaluation framework incorporating confidence-accuracy alignment analysis, Expected Calibration Error (ECE), and Brier Score. Results reveal that existing methods achieve only limited calibration performance; furthermore, task type, prompt engineering, and output format significantly influence uncertainty estimation quality. To foster reproducible research, we open-source our evaluation protocol and analytical toolchain. This work establishes a rigorous, community-accessible benchmark and methodological foundation for advancing LLM reliability and trustworthy AI.
Large language models commonly exhibit overconfidence and poor calibration, and existing calibration methods are often vulnerable to uninformative strategies—such as naive baseline predictions—struggling to balance reliability with utility. This work proposes ACUTE, a protocol that leverages internal activation signals of language models to deliver a general-purpose, sample-efficient, and computationally lightweight approach to confidence estimation. ACUTE is applicable across diverse tasks including multiple-choice question answering, tool calling, and scientific summarization, and introduces EURO, a novel metric that jointly evaluates calibration quality and informativeness. Evaluated on six large models spanning four model families, ACUTE consistently outperforms strong baselines, achieving substantially lower calibration error while enhancing both the trustworthiness and practical utility of model predictions.
This work addresses the calibration problem in probabilistic forecasting: rigorously defining, evaluating, and quantifying the discrepancy between predicted probabilities and the true data-generating distribution to support reliable downstream decision-making. We propose a unified “indistinguishability” framework, formalizing calibration as the extent to which the predicted and true distributions are indistinguishable under a specified class of discriminators. This is the first systematic unification of mainstream calibration metrics—including Expected Calibration Error (ECE) and Kernel Calibration Error (KCE)—as instances of discrimination failure under varying discriminator capacities. Leveraging statistical hypothesis testing, probability theory, and learning theory, we develop computationally tractable estimators for calibration error and establish theoretical links between calibration error and decision-theoretic risk. The framework provides a novel paradigm for calibration analysis and reveals the operational limits of existing metrics in real-world decision contexts.
To address the challenge of ensuring output reliability of large language models (LLMs) in risk-sensitive open-domain question answering, this paper proposes SAFER—a selective abstinence-aware framework. First, it introduces an abstention-aware two-stage sampling mechanism that dynamically adjusts sample size under a finite budget. Second, it filters candidate answers using Clopper–Pearson confidence intervals and conformal risk control thresholds, explicitly bounding both error coverage and the risk of erroneously rejecting correct answers. SAFER is the first method to integrate principled abstention with dual-risk control into selective conformal prediction, enabling task-adaptive acceptance criteria and decoupled calibration–testing. Evaluated on multiple open-domain QA benchmarks, SAFER significantly improves output reliability under strict risk constraints, demonstrating strong robustness, high data efficiency, and flexible risk configurability—even with low sampling budgets.
Existing confidence calibration methods predominantly rely on statistical fitting, neglecting the underlying prior distribution governing calibration curves. This work proposes a novel calibration framework based on the Binomial Process (BPM), the first to model calibration data as a binomial process. We theoretically establish its Lipschitz continuity and high sample efficiency—requiring only $3/B$ samples compared to histogram-based methods (where $B$ is the number of bins). Our approach jointly incorporates prior knowledge and empirical observations, fitting a continuous calibration curve via maximum likelihood estimation and joint optimization. We further introduce the Total Calibration Error (TCE$_{ ext{pm}}$), a consistent and unbiased metric for calibration error assessment. Extensive experiments on both synthetic and real-world datasets demonstrate that our method significantly outperforms state-of-the-art approaches in calibration accuracy, robustness under limited samples, and consistency of error estimation.
Existing Process Reward Models (PRMs) suffer from severe miscalibration of uncertainty, systematically overestimating the success probability of reasoning trajectories—leading to inefficient resource allocation. This work introduces, for the first time, uncertainty calibration into PRMs, proposing a quantile regression–based calibration method and an Instance-Adaptive Scaling (IAS) framework: IAS dynamically adjusts the rollout depth of each reasoning trajectory based on its calibrated confidence score. The approach enables on-demand inference, substantially reducing computational cost while preserving answer accuracy. Experiments demonstrate that our method significantly outperforms baselines in calibration error metrics (e.g., Expected Calibration Error), achieving a 32% average reduction in inference cost on mathematical reasoning tasks without sacrificing accuracy. Our core contributions are: (1) the first uncertainty calibration paradigm specifically designed for PRMs; and (2) a learnable, instance-aware adaptive inference scheduling mechanism.
论文探讨了大语言模型内部因果声明测量方法的失效问题,并提出通过校准、验证等四种方法来解决这些问题。
本文针对条件均值的校准点预测问题,利用共形预测开发了校准置信区间方法,提供不确定性量化,并通过模拟和保险数据集验证了方法的有效性。
论文提出Evidence-Calibration-Stability框架,通过区分证据、校准和稳定性来解决模型不确定性下的假设检验问题。
This study addresses the long-standing stagnation of the sequential calibration error exponent at $O(T^{2/3})$ without explicit constants, which has hindered theoretical progress for over two decades. To overcome this limitation, this work proposes a two-stage recursive labeling strategy combined with an optimization reduction approach, thereby refining the equivalence between sign-preserving games and calibration. By integrating logarithmic instance reductions with explicit parameter composition techniques, it achieves a fundamental theoretical breakthrough. The primary contribution is the first derivation of a calibration error exponent strictly below $2/3$, establishing a new bound of $O(T^{0.662942288})$. This result significantly advances the theory of sequential online calibration by resolving a key open problem that has persisted for more than twenty years.
本文提出一种精确测量大语言模型优化对输出质量影响的方法,通过校准的LLM评分系统,对比不同优化技术在相同提示下的表现。