distribution-aligned calibration

Designs and evaluates procedures that calibrate probabilistic model outputs by aligning predictive-score distributions across models, groups, or datasets. This includes building reweighting or resampling schemes to equalize score distributions and measuring calibration conditional on those matched score distributions.

distribution-alignedcalibration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Modern statistical models are growing increasingly complex in an effort to realistically capture system dynamics. Using standard simulation-based inference, these models may be computationally prohibitive, necessitating the use of model calibration methods. Bayesian score calibration is a computationally efficient framework for model calibration with strong theoretical guarantees. This framework learns an appropriate correction for an approximate model using a small number of simulations from the data-generating process. Currently, only a location-scale transformation has been explored, which may lack the flexibility to correct the complex error introduced by some approximate models. In this paper, we develop two flexible transformations for use in the Bayesian score calibration framework. The first is a polynomial extension, which can appropriately adjust approximate models with location-varying error. The second is a sequential application of Bayesian score calibration, which can accommodate approximate models with posteriors that have low support for the true parameter values. We also discuss an additional diagnostic for use with this framework. We demonstrate the increased flexibility these two approaches provide over Bayesian score calibration in two illustrative simulation studies.

approximate modelsBayesian score calibrationcomplex error

This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.

calibrationclassificationhierarchical relations

Design-marginal calibration of Gaussian process predictive distributions: Bayesian and conformal approaches

Dec 05, 2025
AP
Aurélien Pion
🏛️ Transvalor S.A. | Univ. Paris-Saclay

This paper addresses the calibration of predictive distributions for Gaussian processes (GPs) under interpolation settings, formally defining μ-coverage and μ-probabilistic calibration via the randomized probability integral transform (RPIT) from a design-marginal perspective. We propose two novel methods: CPS-GP (Conformalized Predictive Smoothing GP), which achieves finite-sample marginal calibration, and BCR-GP (Bayesian-Constrained Residual GP), which yields smooth, sharp, and tail-controlled predictive distributions. Technically, both methods integrate leave-one-out residual standardization, generalized normal distribution modeling, cross-validated residual fitting, and Kolmogorov–Smirnov testing. Experiments demonstrate that CPS-GP and BCR-GP significantly outperform Jackknife+ and full-conformal GP in calibration metrics—including empirical coverage, KS statistic, and integrated absolute error—as well as in accuracy, measured by scaled continuous ranked probability score (CRPS). These advances provide a more reliable foundation for uncertainty quantification in applications such as sequential Bayesian optimization.

Calibrating Gaussian process predictive distributions for interpolationControlling dispersion and tail behavior in sequential design predictionsEnsuring marginal calibration through Bayesian and conformal methods

Calibration through the Lens of Indistinguishability

Sep 02, 2025
PG
Parikshit Gopalan
🏛️ Apple | Northeastern University

This work addresses the calibration problem in probabilistic forecasting: rigorously defining, evaluating, and quantifying the discrepancy between predicted probabilities and the true data-generating distribution to support reliable downstream decision-making. We propose a unified “indistinguishability” framework, formalizing calibration as the extent to which the predicted and true distributions are indistinguishable under a specified class of discriminators. This is the first systematic unification of mainstream calibration metrics—including Expected Calibration Error (ECE) and Kernel Calibration Error (KCE)—as instances of discrimination failure under varying discriminator capacities. Leveraging statistical hypothesis testing, probability theory, and learning theory, we develop computationally tractable estimators for calibration error and establish theoretical links between calibration error and decision-theoretic risk. The framework provides a novel paradigm for calibration analysis and reveals the operational limits of existing metrics in real-world decision contexts.

Assessing calibration as indistinguishability between predicted and real worldsDefining and measuring calibration error in predictorsInterpreting predicted probabilities for discrete outcomes

This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.

calibrationdistributional validationprobabilistic forecasting

Latest Papers

What's happening recently
View more

This study addresses the limitation of existing recalibration methods, which often obscure miscalibration in specific regions such as extreme events. To overcome this, we propose an outcome-conditional recalibration post-processing method that leverages quantile recalibration and conditional distribution scaling to achieve precise correction of arbitrary predictive distributions within user-defined regions. By simultaneously preserving global performance and local reliability, the proposed approach significantly enhances conditional calibration on regression benchmark tasks. Furthermore, when applied to electricity price forecasting, it substantially improves calibration in negative-price regimes with negligible accuracy loss. Overall, this work provides more reliable localized guarantees for probabilistic forecasting.

calibrationextreme eventsoutcome-conditional

This work addresses the challenge of accurately evaluating and comparing candidate conditional distributions using only joint samples drawn from the true data-generating process. The authors propose the MIRA score, a novel evaluation metric grounded in the principle that the true and candidate conditional distributions should assign identical probability mass across all regions of the support. By constructing an analytic statistic that directly quantifies consistency between a candidate conditional distribution and the true generative mechanism, MIRA enables unbiased validation without requiring marginal likelihood computation. To the best of the authors’ knowledge, this is the first method to achieve such validation solely from joint samples, while also providing a theoretical reference value and uncertainty estimates. Empirical results on synthetic benchmarks and Bayesian inference tasks demonstrate that MIRA facilitates accurate assessment and reliable model comparison for conditional distributions.

Bayesian inferenceconditional distributiondistribution accuracy

本文提出了一种新的度量方法rankECE,通过比较预测概率相近的点来更准确地估计模型的校准误差,以解决现有ECE估计方法不准确的问题。

CalibrationExpected Calibration Error (ECE)Miscalibration

This work addresses the lack of explicit characterization of the interplay among information, reliability, and uncertainty in existing probabilistic forecast calibration methods. For any proper scoring rule, the authors propose the first general triple-decomposition framework grounded in information algebra and conditional entropy theory, which rigorously decomposes predictive loss into three distinct components: reliability (calibration error), information loss, and irreducible uncertainty. This framework uniquely quantifies the information loss incurred when mapping features to predictive scores and provides a unified interpretation of post-hoc calibration, model ensembling, and boosting strategies. In classification tasks, the approach is successfully applied to calibration evaluation, model aggregation, and staged training, clearly disentangling each component’s contribution to overall predictive uncertainty.

calibrationinformation lossprobabilistic prediction

This work addresses the challenge that probabilistic programs generated by language models often suffer from statistical misspecifications—such as incorrect likelihoods, priors, or parameterizations—that are difficult to detect with conventional unit tests. The paper introduces, for the first time, Bayesian calibration as a central criterion for assessing the correctness of probabilistic programs and proposes a fully unsupervised, reference-free framework for their detection and repair. By integrating Bayesian validation techniques—including posterior predictive checks, simulation-based calibration (SBC), sampling diagnostics (e.g., $\hat{R}$, divergences, effective sample size), and held-out predictive log density—the method generates feedback signals to drive an iterative repair loop within large language models. Evaluated on 200 instances, the approach achieves detection AUCs of 0.97 with reference programs and 62–78% without, substantially outperforming unit testing; repair success rates reach 92% and 100% using GPT-5.1 and Claude, respectively.

Bayesian workflowcalibrationlanguage models

Hot Scholars

CJ

Changwoo J. Lee

Postdoctoral Associate, Duke University
Probabilistic machine learningBayesian statisticsEnvironmental epidemiologyClustering
JJ

Junyi Jessy Li

Associate Professor, The University of Texas at Austin
Computational LinguisticsNatural Language Processing
JJ

Jiaojiao Jiang

The University of New South Wales
Social Network Analysis and Service Virtualisation
ZY

Zohar Yakhini

Faculty Member, Computer Science at IDC Herzeliya
Computational Biology