Score
Designs and implements nonlinear, often monotone, mappings that transform continuous model outputs (e.g., means, variances, or raw scores) into calibrated probabilities or categorical probability assignments. This includes selecting mapping families (for example binning or parametric monotone transforms), fitting the mapping parameters, and evaluating the calibration and reliability of the resulting predictive probabilities.
This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.
Deep neural networks often produce miscalibrated, overconfident probability estimates; existing post-hoc calibration methods struggle to simultaneously preserve instance-level monotonicity—i.e., the original ranking of class probabilities—and expressive power. This paper proposes a novel class of constraint-based, instance-level monotonic calibration methods. We introduce the first linearly parameterized monotonic calibration mapping and formulate calibration as a convex-order-constrained optimization problem, thereby guaranteeing strict monotonicity. Our approach enhances expressivity, interpretability, and robustness without modifying model architecture or requiring retraining, and is applicable to multi-class settings. Extensive experiments across diverse datasets and state-of-the-art models demonstrate that our method consistently outperforms existing SOTA calibration techniques, achieving over 30% average improvement in Expected Calibration Error (ECE) and other calibration metrics. It further exhibits high data efficiency and low computational overhead.
This study addresses the stability of the solution operator with respect to perturbations in the input parameter distribution within the framework of nonparametric Bayesian computer model calibration. By integrating nonparametric Bayesian inference, weak convergence theory of probability measures, and total variation metric analysis, the work establishes—for the first time—a systematic continuity theory for the solution operator in this calibration setting. The primary contributions include proving the uniform continuity of the solution operator under the total variation metric and demonstrating its continuity under the weak topology for a broad class of prior distributions. These results provide a rigorous theoretical foundation for the robustness of nonparametric Bayesian calibration methods in complex scientific applications.
This paper addresses the conceptual ambiguity, incomparability, and heterogeneous objectives plaguing calibration in predictive systems. We propose a unified distributional calibration framework. Methodologically, we introduce— for the first time—two semantic motivations: predictor self-realization (Γ-calibration) and decision-oriented accurate loss estimation; we then construct a formal semantic map leveraging properties of outcome distributions (Γ), swap regret, omniprediction, and multi-granularity grouping generalization. Theoretical contributions include: (i) proving Γ-calibration is equivalent to a specific swap-regret condition; (ii) unifying binary and high-dimensional calibration definitions; (iii) revealing the fundamental role of grouping in both calibration paradigms; and (iv) establishing deep connections to multicalibration and actuarial fairness. Our work provides the first systematic, interpretable, and designable theoretical foundation for calibration in trustworthy AI.
This paper addresses the calibration of predictive distributions for Gaussian processes (GPs) under interpolation settings, formally defining μ-coverage and μ-probabilistic calibration via the randomized probability integral transform (RPIT) from a design-marginal perspective. We propose two novel methods: CPS-GP (Conformalized Predictive Smoothing GP), which achieves finite-sample marginal calibration, and BCR-GP (Bayesian-Constrained Residual GP), which yields smooth, sharp, and tail-controlled predictive distributions. Technically, both methods integrate leave-one-out residual standardization, generalized normal distribution modeling, cross-validated residual fitting, and Kolmogorov–Smirnov testing. Experiments demonstrate that CPS-GP and BCR-GP significantly outperform Jackknife+ and full-conformal GP in calibration metrics—including empirical coverage, KS statistic, and integrated absolute error—as well as in accuracy, measured by scaled continuous ranked probability score (CRPS). These advances provide a more reliable foundation for uncertainty quantification in applications such as sequential Bayesian optimization.
Modern statistical models are growing increasingly complex in an effort to realistically capture system dynamics. Using standard simulation-based inference, these models may be computationally prohibitive, necessitating the use of model calibration methods. Bayesian score calibration is a computationally efficient framework for model calibration with strong theoretical guarantees. This framework learns an appropriate correction for an approximate model using a small number of simulations from the data-generating process. Currently, only a location-scale transformation has been explored, which may lack the flexibility to correct the complex error introduced by some approximate models. In this paper, we develop two flexible transformations for use in the Bayesian score calibration framework. The first is a polynomial extension, which can appropriately adjust approximate models with location-varying error. The second is a sequential application of Bayesian score calibration, which can accommodate approximate models with posteriors that have low support for the true parameter values. We also discuss an additional diagnostic for use with this framework. We demonstrate the increased flexibility these two approaches provide over Bayesian score calibration in two illustrative simulation studies.
This work addresses key challenges in post-hoc calibration—namely nonlinear miscalibration, poor scalability to large numbers of classes, and perturbation of original predictions—by proposing Invertible Logit Transformation (InvLT). InvLT applies a shared-parameter scalar MLP element-wise to pre-softmax logits and incorporates a paired inverse network with soft monotonicity constraints. This design achieves high expressiveness and strong class scalability without introducing class-dependent parameters or requiring model retraining, while rigorously preserving the original classification accuracy. Extensive experiments across diverse image classification benchmarks and model architectures demonstrate that InvLT consistently outperforms existing calibration methods on standard calibration metrics, all while maintaining the original predictive performance without degradation.
本文针对条件均值的校准点预测问题,利用共形预测开发了校准置信区间方法,提供不确定性量化,并通过模拟和保险数据集验证了方法的有效性。
This study addresses the limitation of existing recalibration methods, which often obscure miscalibration in specific regions such as extreme events. To overcome this, we propose an outcome-conditional recalibration post-processing method that leverages quantile recalibration and conditional distribution scaling to achieve precise correction of arbitrary predictive distributions within user-defined regions. By simultaneously preserving global performance and local reliability, the proposed approach significantly enhances conditional calibration on regression benchmark tasks. Furthermore, when applied to electricity price forecasting, it substantially improves calibration in negative-price regimes with negligible accuracy loss. Overall, this work provides more reliable localized guarantees for probabilistic forecasting.
本文通过将概率视为预测方法的输出,统一了不同概率解释,并提出当满足有限校准标准时,可以预测效用分布以指导决策。