Score
Design and execute simulation-based calibration (SBC) workflows that generate data from a model, run inference to obtain posterior samples, and aggregate rank-based diagnostics across many simulated datasets; analyze those diagnostics to detect estimation bias, miscalibration, sampler or software errors, and to quantify the reliability of posterior inference.
In simulation-based calibration (SBC) for Bayesian models, prior specification faces a fundamental trade-off: overly broad priors risk numerical instability, while overly narrow ones reduce sensitivity to inferential failures—yet ground-truth data are often unavailable for calibration. Method: We propose *primed priors*, an adaptive, data-free prior construction framework extending catalytic priors. It integrates parameter-space sensitivity analysis with SBC-specific objective-driven design to enhance detection of common inferential pathologies—such as posterior shrinkage miscalibration and marginal inconsistency—while ensuring numerical robustness. Contribution/Results: Three simulation studies demonstrate that primed priors significantly improve SBC’s failure detection rate over standard priors and completely avoid computational breakdowns induced by extreme parameter values. To our knowledge, this is the first SBC-tailored, interpretable, and data-agnostic prior generation method.
Approximate Bayesian inference often underestimates true uncertainty due to posterior credible intervals that are excessively narrow. This work proposes two simulation-based calibration (SBC)-driven methods for recalibrating approximate posteriors, systematically leveraging the SBC framework to adjust the width of posterior uncertainty intervals and achieve marginal calibration. The approach is applicable to complex model structures, including hierarchical models, and demonstrates consistent efficacy across diverse experimental settings by meaningfully widening posterior intervals. As a result, the proposed recalibration substantially enhances the calibration accuracy and reliability of approximate Bayesian inference.
Likelihood-free and gradient-free parameter calibration in black-box simulators poses significant challenges for Bayesian inference. Method: This paper introduces the first simulation-based, fully amortized, gradient-free, and parallelizable neural Bayesian inference framework, accompanied by the open-source PyTorch package SBI. The framework unifies neural posterior estimation (NPE), neural likelihood estimation (NLE), neural ratio estimation (NRE), and mixture density networks (MDNs), integrating Monte Carlo sampling, Bayesian optimization, and simulation scheduling into a modular, end-to-end workflow with production-ready defaults and comprehensive diagnostic tools. Contribution/Results: Evaluated across physics, biology, and astronomy, SBI substantially lowers the barrier to simulation-based inference, accelerates posterior estimation by multiple-fold, and achieves state-of-the-art reusability and scalability.
This study addresses the lack of reliable validation mechanisms for Bayes factor computations. To this end, we propose two novel calibration methods: (1) an enhanced simulation-based calibration (SBC) procedure, specifically adapted to handle posterior distributions under improper priors; and (2) a binary-prediction-based calibration metric. We comparatively evaluate these against established approaches—including data-averaged posterior checks and the Good test—and find that binary-prediction calibration achieves higher sensitivity under limited computational budgets, whereas SBC detects a broader spectrum of inferential errors. Empirical experiments demonstrate that mainstream R packages—such as *bridgesampling* and *BayesFactor*—exhibit robust performance under default settings. We recommend that new implementations conduct at least several hundred simulation-based calibration runs for rigorous validation. Overall, this work establishes a more efficient and robust framework for validating Bayesian inference, particularly in Bayes factor computation.
This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.
本文提出一种基于三向假设检验的框架,用于量化模拟基础推理中的认知校准不确定性,并评估必要的模拟预算。
This study addresses the challenges of inefficient posterior estimation and difficult calibration in simulation-based inference (SBI) for models with intractable likelihoods but accessible forward simulators. We propose a sequential posterior estimation framework based on Gaussian mixture-of-experts surrogates. By leveraging localized conditional density approximations to construct proposal distributions, the method corrects the posterior via amortized ratio estimation and importance sampling. Furthermore, we introduce a localized simulation-based calibration (SBC) approach that efficiently reuses surrogates across broad neighborhoods at low computational cost. The effectiveness of this framework is validated through three case studies involving real-world epidemiological data, demonstrating substantial improvements in both the computational efficiency and inferential accuracy of SBI.
This work addresses inverse problems in science and engineering—such as parameter inference and detector response unfolding—by proposing a unified simulation-based inference (SBI) framework that systematically integrates Bayesian and frequentist perspectives. Leveraging machine learning techniques, including neural posterior estimation and neural likelihood estimation, the framework enables efficient and general-purpose parameter inference, with extensions to empirical Bayes and unfolding tasks. The paper provides a comprehensive review of SBI methodologies and their application paradigms, while also offering a thorough analysis of validation strategies and inherent limitations. By clarifying best practices and pitfalls, this study advances the reliable deployment and innovative application of SBI in scientific domains.
Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.
This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.