Score
Estimating, calibrating, and propagating uncertainty or confidence in model predictions using probabilistic inference and calibration techniques; used to assess reliability on small datasets, avoid hallucinations in generated designs, and study scaling of stabilized quantiles.
This work addresses the challenge of conducting calibrated Bayesian inference for parametric models whose likelihood functions are intractable, numerically unstable, or computationally prohibitive. Existing approaches lack finite-sample calibration guarantees under such conditions. The authors propose a fully probabilistic inference framework that requires neither a prior nor a likelihood, relying solely on the model’s simulation capability. By leveraging permutation-invariant functions—such as depth functions—to rank parameters and introducing a closed-form rescaling procedure, the method achieves finite-sample frequentist calibration. To the best of the authors’ knowledge, this is the first approach to provide theoretical calibration guarantees in a setting devoid of both likelihood and prior specifications. Empirical evaluations across four benchmark tasks—including differential privacy and the Ising model—as well as a spatial analysis of the 2025 U.S. measles outbreak demonstrate the method’s strong practical utility and robustness.
Bayesian inference in small-sample disciplines (e.g., psychology) often suffers from long-term inferential failure due to misspecified priors, while the true data-generating mechanism remains unknown. Method: We propose a novel framework for calibrating Bayesian credible regions by embedding frequentist validity—specifically, nominal coverage preservation—into Bayesian inference. Our approach jointly leverages Bayesian modeling and a frequency-based calibration objective, solved via stochastic approximation to determine calibrated thresholds; systematic Monte Carlo experiments validate its performance. Contribution/Results: Uncalibrated Bayesian methods frequently yield overly permissive intervals with subnominal coverage. In contrast, our calibrated approach robustly maintains nominal coverage (e.g., 95%) across diverse data-generating mechanisms—including misspecified and nonstandard settings—without requiring knowledge of the true parameter-generating process. This substantially enhances the reliability and reproducibility of Bayesian inference in small-sample contexts.
In statistical inference, confidence sets—especially under complex models or small sample sizes—often fail to achieve nominal coverage levels, particularly in likelihood-free inference (LFI) settings. To address this, we propose TRUST and TRUST++, two distribution-free, simulation-based calibration methods that adapt conformal prediction principles to confidence set construction with redundant parameters, thereby establishing the first distribution-agnostic calibration framework for statistical inference. Our methods guarantee finite-sample local coverage and asymptotic conditional coverage, while enabling self-assessment of simulation cost. Theoretically, we prove their robustness against model misspecification and simulation imperfection. Empirically, TRUST and TRUST++ significantly improve coverage accuracy across both tractable and intractable likelihood models, consistently outperforming existing approaches—especially in small-sample regimes.
This work addresses the issue of overconfidence in statistical inference arising from machine learning approximations in scientific simulations, which can compromise result reliability. To mitigate this, the authors propose two complementary approaches: first, a “balanced” regularization strategy that explicitly suppresses model overconfidence; and second, a simulation-aware Bayesian neural network prior that naturally alleviates overconfidence without additional regularization, even in small-sample regimes. By integrating neural ratio estimation with uncertainty quantification techniques, the proposed methods significantly improve inference calibration, yielding posterior estimates that are either closer to the ground truth or conservatively biased. This enhanced calibration strengthens the credibility of simulation-based inference in scientific applications.
Existing conformal prediction (CP) methods provide finite-sample validity guarantees but lack the ability to quantify support strength for arbitrary events and do not support posterior inference. Method: We propose model-free generalized fiducial inference (MF-GFI), the first framework unifying desiderata-based inference with confidence set approximation via optimal probability measures that approximate fuzzy belief/likelihood pairs, with reliability and accountability as core inferential principles. The method strictly controls Type-I error in finite samples and enables both exact and approximate probabilistic reasoning. We develop a computationally tractable probability approximation algorithm yielding prediction sets with rigorous error guarantees. Contribution/Results: This work establishes the first model-free, accountable, and broadly applicable statistical framework for uncertainty quantification in machine learning—offering finite-sample validity, event-specific support assessment, and principled posterior-like inference without distributional assumptions.
Approximate Bayesian inference often underestimates true uncertainty due to posterior credible intervals that are excessively narrow. This work proposes two simulation-based calibration (SBC)-driven methods for recalibrating approximate posteriors, systematically leveraging the SBC framework to adjust the width of posterior uncertainty intervals and achieve marginal calibration. The approach is applicable to complex model structures, including hierarchical models, and demonstrates consistent efficacy across diverse experimental settings by meaningfully widening posterior intervals. As a result, the proposed recalibration substantially enhances the calibration accuracy and reliability of approximate Bayesian inference.
This study addresses the long-standing challenge of calibrating local false discovery rates (lfdr) in multiple hypothesis testing when true labels are unavailable. The authors introduce, for the first time, a pseudo-labeling mechanism based on spacings between ordered p-values, reframing lfdr calibration as an unsupervised regression problem and thereby enabling the application of classical probability calibration tools. By integrating an empirical Bayes framework with posterior calibration techniques, the method reveals that the widely used q-value approach suffers from substantial miscalibration. Extensive empirical analyses in psychology and neuroscience demonstrate that the proposed approach significantly enhances the reliability and interpretability of lfdr estimates, underscoring its necessity and superiority over existing practices.
This study addresses the challenge of jointly modeling calibration and control parameters in computer model calibration, where the distribution of calibration parameters is unknown while that of control parameters is known. To tackle this issue, the authors propose a nonparametric Bayesian calibration method based on measure decomposition. The approach preserves the known marginal distribution of the control parameters while employing stochastic process modeling and Bayesian inference to construct a posterior distribution over the input space that aligns with field observations. Notably, this work is the first within a nonparametric calibration framework to explicitly maintain the prior distributional properties of the control parameters, thereby substantially enhancing the physical consistency and scientific credibility of the calibration results.
This work addresses a critical limitation in existing model calibration methods, which assess only the reliability of predicted probabilities and fail to evaluate whether the estimated epistemic uncertainty itself is trustworthy—particularly in second-order classification tasks. To bridge this gap, the paper introduces the notion of *cognitive calibration*, a stronger criterion than classical calibration, which measures whether a model’s reported epistemic uncertainty faithfully reflects the dispersion of its predictive distribution around the true label. The authors formalize an evaluation framework and propose the Expected Epistemic Calibration Error (EECE) as a consistent estimator of the true cognitive calibration error (TECE). This reveals failure modes invisible to conventional metrics and leads to an impossibility theorem. Empirical results demonstrate that cognitive calibration provides a coherent and meaningful evaluation standard, under which different uncertainty quantification methods exhibit markedly distinct behaviors despite similar predictive performance.
This work rigorously establishes that predictive Bayesian inference (PBI) suffers from severely miscalibrated posterior uncertainty in practice when the forward predictive model is misspecified, potentially yielding credible sets with coverage approaching zero. We prove for the first time that the posterior concentrates precisely on the target dictated solely by the chosen forward predictive model, and that calibration of inference is guaranteed only when this model fully encompasses the true data-generating mechanism. By integrating predictive recursion algorithms, Bayesian nonparametric theory, and posterior concentration analysis, we demonstrate the fundamental role of the predictive model in determining inferential reliability and delineate the necessary conditions under which PBI can deliver valid uncertainty quantification.