Score
Designs and implements probabilistic models that encode uncertainty over parameters and latent variables using priors and likelihoods, including hierarchical (multilevel) structures; specifies the joint and posterior distributions and derives estimators and summaries. Builds and analyzes Bayesian inference workflows — selecting priors, performing posterior estimation (e.g., MCMC, variational inference, or analytic approximations), and assessing model fit and predictive performance.
Gaussian processes (GPs) in surrogate modeling are highly sensitive to misspecification of covariance hyperparameters—particularly the length-scale parameter θ. While fully Bayesian hierarchical inference improves robustness and uncertainty quantification, its performance critically depends on the choice of prior distributions and Markov Chain Monte Carlo (MCMC) proposal mechanisms—a dependency lacking systematic evaluation in prior work. This paper conducts the first comprehensive study of how alternative priors for θ (uniform, Gamma, inverse-Gamma) and their corresponding MCMC proposals affect posterior sampling efficiency, convergence speed, and predictive performance. Leveraging both synthetic and real-world benchmarks under Bayesian GP inference, we demonstrate that principled alignment between prior and proposal distributions significantly enhances prediction accuracy, improves uncertainty calibration, and accelerates MCMC convergence. Our empirical findings provide actionable guidelines and practical design principles for hyperparameter prior selection in Bayesian GP modeling.
Bayesian hierarchical linear models face challenges including weak between-group separation and computationally expensive MCMC inference in high-dimensional or large-sample settings. Method: This paper systematically compares variational inference (VI), stochastic variational inference (SVI), and MCMC across three canonical hierarchical model classes, using both simulation studies and real-data analyses. Contribution/Results: It provides the first quantitative assessment of VI/SVI versus MCMC in terms of posterior dependency fidelity, accuracy in recovering global effects and cluster structure, and stability of WAIC/DIC. Results show that VI/SVI yield accurate estimates of global regression coefficients and group-level structure at substantially lower computational cost, but sacrifice precision in posterior covariance modeling under weak separation—leading to instability in information criteria. Based on these findings, the study delineates the practical applicability boundary of VI as a computationally efficient alternative to MCMC and offers theoretical grounding and empirical guidance for extending VI to generalized hierarchical models.
To address the high computational cost arising from joint propagation of aleatoric and epistemic uncertainties in Bayesian two-stage inference, this paper proposes an efficient uncertainty propagation framework. First, a representative subset is selected via Pareto-smoothed importance sampling to reduce sampling redundancy. Second, an importance-weighted moment-matching strategy is introduced for lightweight posterior approximation. Third, an iterative mixture-distribution expansion mechanism is developed to jointly model both uncertainty types within surrogate modeling and MICE-based multiple imputation. The method preserves posterior accuracy while significantly reducing the cost of multi-model fitting—achieving several-fold improvements in computational efficiency. It establishes a scalable paradigm for Bayesian inference under complex, heterogeneous uncertainty scenarios.
Gibbs sampling for Bayesian mixture models suffers from slow mixing in the marginal posterior over component assignments and struggles to jointly perform model selection and parameter inference. Method: We propose two novel joint-sampling MCMC algorithms: (1) a collapsed Gibbs sampler incorporating unconventional move sets, and (2) a prior-driven, rejection-free component allocation sampler. Both methods jointly update observation assignments and the number of components, unifying model fitting and dimensionality inference. Contribution/Results: Our approaches eliminate the need for post-hoc model selection and substantially improve Markov chain mixing efficiency. In latent class analysis tasks, they reduce mixing time by several-fold compared to state-of-the-art methods while achieving comparable or superior posterior inference accuracy. The framework provides an efficient, fully automated computational solution for high-dimensional Bayesian nonparametric modeling.
Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.
This work addresses the lack of explicit confidence modeling for various sources of uncertainty in Bayesian inference by proposing a general extension framework that, for the first time, explicitly incorporates confidence in key uncertainty components—such as the prior and likelihood—into Bayesian modeling. The framework not only introduces a novel regularization mechanism but also provides a unified approach to inducing model sparsity. Without compromising theoretical rigor, the method achieves controllable sparsity across diverse models, including linear regression, logistic regression, and Bayesian neural networks, thereby significantly enhancing both interpretability and generalization performance.
Bayesian inference remains challenging for statisticians and learners due to conceptual ambiguities in its philosophical foundations, difficulties in prior specification, and computational complexity. Method: This paper provides a rigorous yet accessible pedagogical framework for Bayesian inference, systematically integrating core components—including Bayes’ theorem, prior modeling, posterior inference, Bayesian hypothesis testing via Bayes factors, and predictive analysis—while clarifying fundamental distinctions from frequentist paradigms in identifiability, asymptotic theory, and decision-theoretic concepts (e.g., loss functions, credible intervals). It bridges analytical derivations with modern simulation techniques such as MCMC, using canonical statistical models as unifying exemplars. Contribution/Results: The framework innovatively connects foundational concepts to advanced topics—including hierarchical modeling, nonparametric Bayesian methods, and spatiotemporal analysis—and has been successfully applied in political science, network analysis, and spatial statistics, substantially lowering the barrier to learning and applying Bayesian methods in practice.
This work proposes a probabilistic inference framework that integrates inductive biases to address the challenges of uncertainty quantification in deep sequential models. While traditional Bayesian approaches struggle with prior specification and inference accuracy in large-scale networks, the proposed method establishes a theoretical connection between Transformer attention mechanisms and sparse Gaussian processes, enabling scalable approximate Bayesian inference. It introduces cross-domain inducing points derived from HiPPO operators to support long-range historical modeling in online learning settings. Furthermore, self-supervised signals are leveraged to enrich the probabilistic structure of latent variables in sequence generation. The resulting approach significantly enhances the uncertainty quantification capability, probabilistic expressiveness, and scalability of deep sequential models, all while maintaining competitive predictive performance.
This work addresses the high implementation complexity and accessibility barriers of inference algorithms in Bayesian nonparametric modeling by proposing a flexible Dirichlet process (DP) framework implemented in R. The framework encapsulates the DP as a reusable object that supports density estimation, clustering, and hierarchical model prior construction, while automatically performing Markov chain Monte Carlo (MCMC) posterior inference. Users can either directly apply pre-specified models or customize base distributions and mixture structures without manually implementing sampling algorithms. By abstracting away computational intricacies while preserving substantial modeling flexibility, this approach significantly lowers the practical barrier to applying Bayesian nonparametric methods across a wide range of statistical analysis tasks.