diagnose posterior distributions

Designs, implements, and interprets diagnostic procedures and visualizations for posterior distributions produced by Bayesian models and approximations, including checks of calibration, coverage, multimodality, and model misspecification. Uses comparisons between approximate and reference posteriors and posterior predictive checks to quantify deficiencies and guide model revision or re‑estimation.

diagnoseposteriordistributions

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Approximate Bayesian inference often underestimates true uncertainty due to posterior credible intervals that are excessively narrow. This work proposes two simulation-based calibration (SBC)-driven methods for recalibrating approximate posteriors, systematically leveraging the SBC framework to adjust the width of posterior uncertainty intervals and achieve marginal calibration. The approach is applicable to complex model structures, including hierarchical models, and demonstrates consistent efficacy across diverse experimental settings by meaningfully widening posterior intervals. As a result, the proposed recalibration substantially enhances the calibration accuracy and reliability of approximate Bayesian inference.

approximate posteriorBayesian inferenceposterior recalibration

Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.

Bayesian decision proceduresexperimental designgeneralized posteriors

Bayesian model criticism using uniform parametrization checks

Mar 24, 2025
CT
Christian T. Covington
🏛️ Harvard T.H. Chan School of Public Health

This paper addresses the lack of directionality, statistical rigor, and computational efficiency in Bayesian model misspecification diagnostics. We propose a novel diagnostic framework based on Uniform Parameterization Checks (UPCs). Our core innovation is the first systematic exploitation of the theoretical property that posterior samples should follow the prior distribution under correct model specification; leveraging probability integral transforms, UPCs uniformly map all random components—across prior, likelihood, and data—into independent *u*-values. This enables interpretable, aggregable, statistically rigorous, and computationally efficient diagnostics (≈ cost of one posterior sample) for individual model components (prior, likelihood, data subsets). UPCs support targeted detection of specific misspecification types—including dependence structure violations, tail discrepancies, and missing correlations—and accurately localize sources of misspecification in both synthetic and real-world examples. Theoretical analysis establishes consistency of the proposed tests.

Detect model misspecification in Bayesian analysisIdentify incorrect aspects of likelihood or priorUse uniform reparametrization for rigorous hypothesis testing

Bridge Sampling Diagnostics

Aug 20, 2025
GM
Giorgio Micaletto
🏛️ Bocconi University | Aalto University

Bayesian model selection and averaging rely on the marginal likelihood, whose exact computation is intractable for complex models; standard estimators such as bridge sampling often yield high-variance approximations. This paper proposes a diagnostic, low-overhead framework to assess the reliability of marginal likelihood estimates: it introduces Pareto-$hat{k}$ diagnostics and block reordering into the bridge sampling pipeline, enabling robust quantification of Monte Carlo standard error (MCSE) without additional posterior sampling. The method integrates bridge sampling, MCSE estimation, and a dual-diagnostic mechanism. In simulation studies and real-world posterior distributions from posteriordb, it substantially reduces estimator variability and enhances credibility. The resulting tool provides a reproducible, verifiable, and practical solution for Bayesian model comparison.

Assessing bridge sampling variability and estimate reliabilityDiagnosing Monte Carlo standard error without repeated inferenceEstimating marginal likelihood for Bayesian model selection

Sequential Design with Posterior and Posterior Predictive Probabilities

Apr 01, 2025
LH
Luke Hagar
🏛️ McGill University | McGill University Health Centre

In Bayesian sequential trials, error rate evaluation relies on computationally expensive Monte Carlo simulations, hindering efficient optimization of sample size and decision thresholds. Method: This paper establishes, for the first time, analytical functional relationships between posterior and posterior predictive probabilities and sample size. Leveraging Bayesian decision theory and asymptotic analysis—combined with numerical fitting and error-rate inversion—the method enables precise error-rate assessment for any sample size using only two simulations, and rapidly identifies optimal design parameters. Contribution/Results: The approach drastically reduces computational cost while achieving error-rate control accuracy comparable to conventional simulation-based methods. In two real-world case studies, it attains exact error-rate calibration and accelerates design optimization by several orders of magnitude. This provides a scalable, verifiable, and highly efficient design paradigm for Bayesian adaptive trials.

Efficient error rate assessment for Bayesian sequential designsModeling probabilities as functions of sample sizeOptimal sample size determination using posterior probabilities

Latest Papers

What's happening recently
View more

This study addresses the common challenge in medical diagnostic research where critical performance metrics—such as sensitivity and specificity—cannot be directly computed due to missing data in 2×2 contingency tables. The authors propose a hierarchical Bayesian model that systematically handles two representative yet challenging missing-data structures, enabling reliable posterior inference and uncertainty quantification for both missing cell counts and diagnostic accuracy measures, even under weak identifiability conditions. By integrating Markov chain Monte Carlo (MCMC) sampling, the method successfully reconstructs complete contingency tables and accurately estimates diagnostic performance while appropriately characterizing parameter uncertainty, as demonstrated on simulated benchmark data from breast MRI studies.

2x2 tableBayesian inferencediagnostic accuracy

This study addresses the lack of theoretical guarantees regarding the impact of predictive design distributions on inference validity in high-dimensional Bayesian regression. The authors propose a novel parametric martingale posterior approach based on sequential one-step-ahead predictive quantities, which operates without Markov chain Monte Carlo methods and, for the first time, systematically elucidates the critical role of predictive design distributions. The method satisfies weak identifiability and design invariance, and integrates predictive resampling with high-dimensional regularization techniques. In simulation experiments, it demonstrates both computational efficiency and stable performance in Bayesian predictive inference.

Bayesian regressiondesign invariancehigh-dimensional regression

This work addresses the challenge that probabilistic programs generated by language models often suffer from statistical misspecifications—such as incorrect likelihoods, priors, or parameterizations—that are difficult to detect with conventional unit tests. The paper introduces, for the first time, Bayesian calibration as a central criterion for assessing the correctness of probabilistic programs and proposes a fully unsupervised, reference-free framework for their detection and repair. By integrating Bayesian validation techniques—including posterior predictive checks, simulation-based calibration (SBC), sampling diagnostics (e.g., $\hat{R}$, divergences, effective sample size), and held-out predictive log density—the method generates feedback signals to drive an iterative repair loop within large language models. Evaluated on 200 instances, the approach achieves detection AUCs of 0.97 with reference programs and 62–78% without, substantially outperforming unit testing; repair success rates reach 92% and 100% using GPT-5.1 and Claude, respectively.

Bayesian workflowcalibrationlanguage models

Hot Scholars

UV

Umberto Villa

Biomedical Engineering and Oden Institute, UT Austin
PhotoacousticUltrasoundImaging ScienceInverse Problems
DX

Dongbin Xiu

Professor of Mathematics, The Ohio State University
applied and computational mathematicsuncertainty quantification