Score
Design, implement, and analyze Bayesian models and inference procedures that estimate prevalence (population class proportions) from observed data, producing posterior distributions and credible intervals for prevalence parameters. Work includes specifying priors and likelihoods (including kernel-density-based likelihoods), performing posterior inference, and integrating or comparing Bayesian prevalence estimates with maximum-likelihood or other quantification methods.
In class distribution estimation (quantification), persistent challenges include training-set prevalence bias and reliable uncertainty quantification—particularly in terms of confidence interval width, coverage accuracy, and calibration. To address these, we propose Precise Quantifier (PQ), a novel Bayesian quantification framework that explicitly models key factors such as classifier discriminative capability, labeled and unlabeled dataset sizes, and label noise. PQ jointly optimizes bias correction and uncertainty calibration through principled posterior inference. Theoretical analysis establishes PQ’s consistency and asymptotic normality under mild assumptions. Extensive experiments across diverse real-world and synthetic benchmarks demonstrate that PQ yields significantly narrower confidence intervals while achieving near-perfect coverage calibration—outperforming classical bootstrap-based approaches and state-of-the-art Bayesian quantifiers. This work introduces a new paradigm for distribution inference that simultaneously delivers high estimation precision and well-calibrated, reliable uncertainty quantification.
Imperfect diagnostic tests—characterized by false positives and false negatives—introduce bias in estimating STI prevalence and associated risk factors. To address this, we systematically compare four misclassification-correction methods and propose, for the first time, a Bayesian internal correction logistic regression model incorporating informative priors. This model jointly estimates test sensitivity, specificity, and true prevalence via MCMC-based inference, enabling coherent parameter optimization. It demonstrates superior performance in low-prevalence, small-sample, and rare-event settings: yielding the narrowest confidence intervals for baseline prevalence, the most stable intercept estimates, and up to a 37% reduction in relative estimation error compared to naïve approaches; it also substantially improves covariate effect estimation. The framework provides a robust, generalizable methodological solution for accurately assessing the true STI burden in resource-limited settings.
Bayesian experimental design with nuisance parameters remains challenging due to the need to account for their prior uncertainty while maintaining statistical rigor and computational feasibility. Method: This paper proposes a fully Bayesian framework that explicitly models prior uncertainty over all parameters—including nuisance parameters—during the design stage. It introduces Bayesian additive regression trees (BART) to the experimental design literature for the first time, integrating asymptotic posterior approximations with Monte Carlo simulation to efficiently optimize sample size and decision rules under both fixed and adaptive designs. Contribution/Results: The approach significantly reduces computational burden compared to conventional resampling-intensive methods, while preserving statistical power and robust operating characteristics. Key innovations include: (1) unified quantification of nuisance parameter uncertainty; (2) a BART-driven design function learning mechanism that enhances interpretability and generalizability; and (3) an end-to-end Bayesian design pipeline balancing robustness, efficiency, and practical implementation.
This study addresses the limited Bayesian statistical background among mathematical epidemiologists by presenting a systematic Bayesian inference framework for the classic SIR model. Starting from the underlying disease transmission dynamics, the authors derive the likelihood function, specify biologically plausible prior distributions for the transmission and recovery rates, and implement posterior sampling using the Metropolis–Hastings algorithm. The work integrates Bayesian inference and Markov chain Monte Carlo (MCMC) methods into a self-contained, accessible tutorial that substantially lowers the technical barrier for newcomers. By offering a reproducible and user-friendly implementation, this approach provides an entry-level paradigm that effectively balances pedagogical clarity with practical applicability for parameter estimation in infectious disease modeling.
Bayesian inference for finite-population surveys is challenging when sampling units exhibit complex dependencies (e.g., spatial, network, or structural) and nonresponse is nonignorable. Method: We propose a unified hierarchical modeling framework that integrates graphical models and spatial random fields to characterize multivariate dependence; formally adopts the “unapologetic Bayesian” paradigm, embedding design-based weights (e.g., Horvitz–Thompson) naturally into prior and likelihood specifications; incorporates causal ignorability analysis to ensure identifiability under missing-not-at-random (MNAR) mechanisms; and employs MCMC and variational inference for scalable computation. Contribution/Results: The framework achieves improved small-area estimation accuracy and more reliable uncertainty quantification in two empirical spatial finite-population analyses. It rigorously reconciles design-based consistency with model-based flexibility, providing theoretical guarantees for valid Bayesian inference under complex survey designs and nonignorable nonresponse.
Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.
This work addresses the high implementation complexity and accessibility barriers of inference algorithms in Bayesian nonparametric modeling by proposing a flexible Dirichlet process (DP) framework implemented in R. The framework encapsulates the DP as a reusable object that supports density estimation, clustering, and hierarchical model prior construction, while automatically performing Markov chain Monte Carlo (MCMC) posterior inference. Users can either directly apply pre-specified models or customize base distributions and mixture structures without manually implementing sampling algorithms. By abstracting away computational intricacies while preserving substantial modeling flexibility, this approach significantly lowers the practical barrier to applying Bayesian nonparametric methods across a wide range of statistical analysis tasks.
This study addresses Bayesian inference for low-dimensional target parameters in semiparametric models, particularly under the presence of complex nuisance components that may compromise frequentist properties. To this end, we construct posterior distributions by integrating estimating function methods with nonparametric Bayesian techniques—such as Dirichlet processes and Bayesian bootstrap—under conditions weaker than the classical stochastic equicontinuity assumption. We establish asymptotic normality and consistency of the resulting posterior, rigorously identifying the key assumptions required to guarantee desirable frequentist behavior. The theoretical analysis systematically elucidates how relaxing these assumptions affects inferential performance. Extensive simulations corroborate the effectiveness of the proposed methodology, demonstrating its robustness and accuracy in practical settings.
This study addresses the inefficiency of Markov chain Monte Carlo (MCMC) sampling in Bayesian models for highly zero-inflated count data, where strong posterior dependencies hinder effective exploration. The work proposes the first integration of marginal data augmentation into a semiparametric Bayesian count regression framework. By introducing working parameters to rescale latent variables associated with zero observations, the method substantially alleviates posterior correlations. This approach markedly improves MCMC mixing and convergence rates, outperforming existing sampling strategies in both synthetic experiments and real-world modeling of subnational mortality counts in Austria. The proposed technique thus offers an efficient and reliable Bayesian inference pathway for high-dimensional zero-inflated count data.