Score
Designs and evaluates procedures that calibrate probabilistic model outputs by aligning predictive-score distributions across models, groups, or datasets. This includes building reweighting or resampling schemes to equalize score distributions and measuring calibration conditional on those matched score distributions.
This paper addresses the misalignment between recommended items and users’ historical interest distributions in recommender systems, offering a systematic survey of recent advances in calibrated recommendation. Methodologically, it integrates statistical matching, constrained optimization, and re-ranking techniques, complemented by empirical evaluation and theoretical modeling to assess improvements in recommendation diversity, bias mitigation, and fairness enhancement. The work establishes the first comprehensive conceptual framework for calibrated recommendation, synthesizes empirical evidence on method effectiveness, and identifies cross-domain challenges—including the utility-calibration trade-off and difficulties in modeling dynamic user interests. Key practical contributions include scalable calibration mechanism design, multi-objective collaborative optimization, and causality-driven fairness assurance. Collectively, these advances bridge the gap between theoretical foundations and real-world deployment, providing a principled methodology for operationalizing calibrated recommendation.
Modern statistical models are growing increasingly complex in an effort to realistically capture system dynamics. Using standard simulation-based inference, these models may be computationally prohibitive, necessitating the use of model calibration methods. Bayesian score calibration is a computationally efficient framework for model calibration with strong theoretical guarantees. This framework learns an appropriate correction for an approximate model using a small number of simulations from the data-generating process. Currently, only a location-scale transformation has been explored, which may lack the flexibility to correct the complex error introduced by some approximate models. In this paper, we develop two flexible transformations for use in the Bayesian score calibration framework. The first is a polynomial extension, which can appropriately adjust approximate models with location-varying error. The second is a sequential application of Bayesian score calibration, which can accommodate approximate models with posteriors that have low support for the true parameter values. We also discuss an additional diagnostic for use with this framework. We demonstrate the increased flexibility these two approaches provide over Bayesian score calibration in two illustrative simulation studies.
This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.
This paper addresses the calibration of predictive distributions for Gaussian processes (GPs) under interpolation settings, formally defining μ-coverage and μ-probabilistic calibration via the randomized probability integral transform (RPIT) from a design-marginal perspective. We propose two novel methods: CPS-GP (Conformalized Predictive Smoothing GP), which achieves finite-sample marginal calibration, and BCR-GP (Bayesian-Constrained Residual GP), which yields smooth, sharp, and tail-controlled predictive distributions. Technically, both methods integrate leave-one-out residual standardization, generalized normal distribution modeling, cross-validated residual fitting, and Kolmogorov–Smirnov testing. Experiments demonstrate that CPS-GP and BCR-GP significantly outperform Jackknife+ and full-conformal GP in calibration metrics—including empirical coverage, KS statistic, and integrated absolute error—as well as in accuracy, measured by scaled continuous ranked probability score (CRPS). These advances provide a more reliable foundation for uncertainty quantification in applications such as sequential Bayesian optimization.
This work addresses the calibration problem in probabilistic forecasting: rigorously defining, evaluating, and quantifying the discrepancy between predicted probabilities and the true data-generating distribution to support reliable downstream decision-making. We propose a unified “indistinguishability” framework, formalizing calibration as the extent to which the predicted and true distributions are indistinguishable under a specified class of discriminators. This is the first systematic unification of mainstream calibration metrics—including Expected Calibration Error (ECE) and Kernel Calibration Error (KCE)—as instances of discrimination failure under varying discriminator capacities. Leveraging statistical hypothesis testing, probability theory, and learning theory, we develop computationally tractable estimators for calibration error and establish theoretical links between calibration error and decision-theoretic risk. The framework provides a novel paradigm for calibration analysis and reveals the operational limits of existing metrics in real-world decision contexts.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
This study addresses the limitation of existing recalibration methods, which often obscure miscalibration in specific regions such as extreme events. To overcome this, we propose an outcome-conditional recalibration post-processing method that leverages quantile recalibration and conditional distribution scaling to achieve precise correction of arbitrary predictive distributions within user-defined regions. By simultaneously preserving global performance and local reliability, the proposed approach significantly enhances conditional calibration on regression benchmark tasks. Furthermore, when applied to electricity price forecasting, it substantially improves calibration in negative-price regimes with negligible accuracy loss. Overall, this work provides more reliable localized guarantees for probabilistic forecasting.
This work addresses the challenge of accurately evaluating and comparing candidate conditional distributions using only joint samples drawn from the true data-generating process. The authors propose the MIRA score, a novel evaluation metric grounded in the principle that the true and candidate conditional distributions should assign identical probability mass across all regions of the support. By constructing an analytic statistic that directly quantifies consistency between a candidate conditional distribution and the true generative mechanism, MIRA enables unbiased validation without requiring marginal likelihood computation. To the best of the authors’ knowledge, this is the first method to achieve such validation solely from joint samples, while also providing a theoretical reference value and uncertainty estimates. Empirical results on synthetic benchmarks and Bayesian inference tasks demonstrate that MIRA facilitates accurate assessment and reliable model comparison for conditional distributions.
本文提出了一种新的度量方法rankECE,通过比较预测概率相近的点来更准确地估计模型的校准误差,以解决现有ECE估计方法不准确的问题。
This work addresses the lack of explicit characterization of the interplay among information, reliability, and uncertainty in existing probabilistic forecast calibration methods. For any proper scoring rule, the authors propose the first general triple-decomposition framework grounded in information algebra and conditional entropy theory, which rigorously decomposes predictive loss into three distinct components: reliability (calibration error), information loss, and irreducible uncertainty. This framework uniquely quantifies the information loss incurred when mapping features to predictive scores and provides a unified interpretation of post-hoc calibration, model ensembling, and boosting strategies. In classification tasks, the approach is successfully applied to calibration evaluation, model aggregation, and staged training, clearly disentangling each component’s contribution to overall predictive uncertainty.
This work addresses the challenge that probabilistic programs generated by language models often suffer from statistical misspecifications—such as incorrect likelihoods, priors, or parameterizations—that are difficult to detect with conventional unit tests. The paper introduces, for the first time, Bayesian calibration as a central criterion for assessing the correctness of probabilistic programs and proposes a fully unsupervised, reference-free framework for their detection and repair. By integrating Bayesian validation techniques—including posterior predictive checks, simulation-based calibration (SBC), sampling diagnostics (e.g., $\hat{R}$, divergences, effective sample size), and held-out predictive log density—the method generates feedback signals to drive an iterative repair loop within large language models. Evaluated on 200 instances, the approach achieves detection AUCs of 0.97 with reference programs and 62–78% without, substantially outperforming unit testing; repair success rates reach 92% and 100% using GPT-5.1 and Claude, respectively.