Score
Designs and implements Bayesian nonparametric mixture models using Dirichlet process priors and empirical-Bayes procedures: constructs DP mixture models to estimate unknown mixing distributions and induced marginal probability mass functions, implements algorithms to obtain DP posteriors and empirical-Bayes estimates, and analyzes their asymptotic and concentration properties (e.g., posterior concentration rates).
To address the challenge of nonparametrically modeling mixed component distributions in heterogeneous data, this paper proposes a finite mixture model with nonparametric components, where each component density is itself modeled via a Dirichlet process mixture (DPM) prior. First, we establish identifiability conditions for the mixing components under this framework. Second, we theoretically prove that the posterior contraction rate for component densities is polynomial—significantly faster than the logarithmic rate typical of conventional deconvolution for mixing measures. Third, to enable efficient Bayesian inference, we design a tailored MCMC algorithm. Extensive simulations and real-data analyses demonstrate the method’s high accuracy and robustness in identifying latent subgroups, estimating population-level and component-specific densities. The approach thus offers both rigorous theoretical guarantees and practical utility for complex heterogeneous data analysis.
This work addresses the high implementation complexity and accessibility barriers of inference algorithms in Bayesian nonparametric modeling by proposing a flexible Dirichlet process (DP) framework implemented in R. The framework encapsulates the DP as a reusable object that supports density estimation, clustering, and hierarchical model prior construction, while automatically performing Markov chain Monte Carlo (MCMC) posterior inference. Users can either directly apply pre-specified models or customize base distributions and mixture structures without manually implementing sampling algorithms. By abstracting away computational intricacies while preserving substantial modeling flexibility, this approach significantly lowers the practical barrier to applying Bayesian nonparametric methods across a wide range of statistical analysis tasks.
Existing Bayesian mixture models assume independence among atoms of random probability measures, neglecting inter-atomic interactions—such as repulsion or attraction—that are critical for meaningful clustering. Method: We propose the first unified framework enabling both repulsive and attractive interactions among atoms, abandoning the classical independence assumption. We derive closed-form expressions for the posterior, marginal, and predictive distributions of interacting-atom random measures; avoid prespecifying finite point process structures; and pioneer the integration of Palm calculus into Bayesian nonparametric inference. Our hierarchical model combines Poisson, Gibbs, and determinantal point processes with shot-noise Cox processes to flexibly encode interaction patterns. Contribution/Results: The framework supports efficient MCMC sampling, ensures prior interpretability and algorithmic convergence, and significantly improves cluster separation and interpretability. Extensive experiments on synthetic and real-world datasets validate its effectiveness and robustness.
Gibbs sampling for Bayesian mixture models suffers from slow mixing in the marginal posterior over component assignments and struggles to jointly perform model selection and parameter inference. Method: We propose two novel joint-sampling MCMC algorithms: (1) a collapsed Gibbs sampler incorporating unconventional move sets, and (2) a prior-driven, rejection-free component allocation sampler. Both methods jointly update observation assignments and the number of components, unifying model fitting and dimensionality inference. Contribution/Results: Our approaches eliminate the need for post-hoc model selection and substantially improve Markov chain mixing efficiency. In latent class analysis tasks, they reduce mixing time by several-fold compared to state-of-the-art methods while achieving comparable or superior posterior inference accuracy. The framework provides an efficient, fully automated computational solution for high-dimensional Bayesian nonparametric modeling.
This paper addresses the longstanding limitation in posterior summarization for nonparametric Bayesian mixture models—where inference has predominantly focused on random partition point estimates, neglecting direct inference on the mixing measure itself. We propose a decision-theoretic framework that prioritizes the mixing measure as the primary inferential target. Methodologically, we introduce, for the first time, a model-agnostic variant of the sliced Wasserstein distance, integrated with generalized geodesic projection and optimization on the symmetric positive-definite matrix manifold; leveraging the linear structure of Gaussian mixing measures, our approach delivers coherent point estimates of the mixing measure, density function, and random partition simultaneously. Compared to conventional paradigms, our method preserves statistical validity under complex dependency structures, substantially improves geometric coherence and computational efficiency in posterior summarization, and unifies support for both density estimation and clustering inference.
This study presents the first systematic investigation into the statistical properties of the Dirichlet process when employed as a sampling distribution and introduces a Bayesian inference framework for its base measure and concentration parameter. Treating the Dirichlet process as a data-generating mechanism, the authors develop a joint inference approach for these two key parameters by integrating Bayesian nonparametric modeling with Markov chain Monte Carlo algorithms, leveraging observed histogram sequences. The proposed methodology is validated through extensive experiments on both synthetic and real-world datasets, demonstrating its effectiveness and practical utility. This work addresses a notable gap in the literature by providing a principled solution to parameter inference in Dirichlet process models, thereby advancing the theoretical and applied understanding of this foundational Bayesian nonparametric construct.
This study addresses the challenge of inferring underlying population distributions from aggregated data available only in the form of histograms or frequency tables. The authors propose a nonparametric Bayesian inference method based on mixture models, employing reversible-jump Markov chain Monte Carlo to fit Gaussian mixtures with either a finite or countably infinite number of components. This work represents the first systematic application of a nonparametric Bayesian framework to histogram data analysis. Furthermore, by leveraging Dirichlet processes, the approach jointly models multiple histograms, enabling information sharing across groups and providing posterior probabilities to quantify homogeneity among them. Empirical evaluations demonstrate that the method effectively reconstructs complex distributions from large-scale aggregated data and offers principled clustering and homogeneity assessment.
This study investigates the posterior asymptotic behavior of Dirichlet process mixture models when the true data-generating distribution is a finite mixture with $K$ components. By integrating the stick-breaking representation, Wasserstein distance, and posterior contraction theory, the work demonstrates that the model adaptively recovers the true number of components $K$ at the parametric $n^{-1/2}$ rate. Moreover, it reveals a phase transition phenomenon: when estimation accuracy exceeds the $n^{-1/2}$ threshold, the required number of mixture components grows on the order of $\log n$. The analysis further establishes that only $O(\log n)$ components are sufficient to replicate the full posterior’s clustering structure, and any truncated model with at least $K$ components simultaneously achieves optimal posterior contraction rates for both density estimation and the mixing measure.
This study addresses the challenge of estimating an unknown mixing distribution in Poisson compound decision problems by systematically comparing a Bayesian empirical Bayes approach based on the Dirichlet process posterior with a quasi-Bayesian method derived from Newton’s algorithm through the lens of g-modeling. It establishes, for the first time, a frequentist convergence theory unifying the two methods, demonstrating that their induced marginal distributions exhibit comparable concentration properties and yield identical regret decay rates. The theoretical analysis is extended to multivariate settings via concentration inequalities and regret bound derivations. Numerical experiments confirm that the quasi-Bayesian method achieves estimation accuracy on par with its Bayesian counterpart while substantially reducing computational cost, with pronounced advantages in high-dimensional scenarios.
This study addresses the challenge of estimating sums of functions over observed and unobserved variables in Poisson mixture models, aiming to overcome limitations in existing nonparametric empirical Bayes theory. To this end, it proposes a nonparametric framework based on regularized and coarsened minimum distance estimation, which achieves near-parametric convergence rates under finite support assumptions. Furthermore, precise minimax regret bounds are derived for two representative classes of summands. Theoretically, the authors establish that the plug-in estimator converges asymptotically to the oracle Bayes estimator, providing finite-sample bounds for both total intensity and hyper-mean that match minimax optimal rates. Extensive experiments on synthetic and real-world datasets validate the effectiveness of the proposed approach.