dirichlet process clustering

Designs and fits Dirichlet-process-based nonparametric mixture models that jointly model many histograms or distributional summaries to cluster populations and share information across them. Builds inference procedures and diagnostics to estimate cluster assignments, component density/histogram parameters, and posterior probabilities of homogeneity among the histograms.

dirichletprocessclustering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the high implementation complexity and accessibility barriers of inference algorithms in Bayesian nonparametric modeling by proposing a flexible Dirichlet process (DP) framework implemented in R. The framework encapsulates the DP as a reusable object that supports density estimation, clustering, and hierarchical model prior construction, while automatically performing Markov chain Monte Carlo (MCMC) posterior inference. Users can either directly apply pre-specified models or customize base distributions and mixture structures without manually implementing sampling algorithms. By abstracting away computational intricacies while preserving substantial modeling flexibility, this approach significantly lowers the practical barrier to applying Bayesian nonparametric methods across a wide range of statistical analysis tasks.

Bayesian nonparametric modelsclusteringdensity estimation

This study addresses the challenge of inferring underlying population distributions from aggregated data available only in the form of histograms or frequency tables. The authors propose a nonparametric Bayesian inference method based on mixture models, employing reversible-jump Markov chain Monte Carlo to fit Gaussian mixtures with either a finite or countably infinite number of components. This work represents the first systematic application of a nonparametric Bayesian framework to histogram data analysis. Furthermore, by leveraging Dirichlet processes, the approach jointly models multiple histograms, enabling information sharing across groups and providing posterior probabilities to quantify homogeneity among them. Empirical evaluations demonstrate that the method effectively reconstructs complex distributions from large-scale aggregated data and offers principled clustering and homogeneity assessment.

Bayesian mixture modelsbinned datahistograms

A Bayesian approach to learning mixtures of nonparametric components

Dec 15, 2025
YZ
Yilei Zhang
🏛️ University of Michigan | University of Texas at Dallas | AT&T Data Science and AI Research

To address the challenge of nonparametrically modeling mixed component distributions in heterogeneous data, this paper proposes a finite mixture model with nonparametric components, where each component density is itself modeled via a Dirichlet process mixture (DPM) prior. First, we establish identifiability conditions for the mixing components under this framework. Second, we theoretically prove that the posterior contraction rate for component densities is polynomial—significantly faster than the logarithmic rate typical of conventional deconvolution for mixing measures. Third, to enable efficient Bayesian inference, we design a tailored MCMC algorithm. Extensive simulations and real-data analyses demonstrate the method’s high accuracy and robustness in identifying latent subgroups, estimating population-level and component-specific densities. The approach thus offers both rigorous theoretical guarantees and practical utility for complex heterogeneous data analysis.

Develops Bayesian nonparametric mixture models for heterogeneous dataEstablishes posterior contraction rates and efficient MCMC inferenceIdentifies conditions for mixture component distribution identifiability

This study addresses the trade-off between computational efficiency and full posterior inference in Bayesian clustering of multivariate binary data by proposing a Bayesian mixture model that integrates a penalized complexity prior with an asymmetric Dirichlet prior. The approach accommodates a large number of latent components while enabling intuitive control over the distribution of the number of clusters through its asymmetric prior structure. Computational feasibility is ensured via an efficient Markov chain Monte Carlo (MCMC) algorithm. Empirical evaluations on both simulated and real-world ecological presence–absence species data demonstrate that the proposed model performs comparably or superiorly to existing methods, successfully achieving a balance among computational efficiency, Bayesian inferential completeness, and interpretability in cluster analysis.

Bayesian inferenceclusteringcomputational efficiency

Identifying, evaluating, and validating clustering structures in Bayesian mixture models has long been hindered by the complexity of the posterior distribution. This work proposes CliPS, a novel approach that reformulates mixture models as point processes and introduces a low-dimensional parametric functional mapping to transform MCMC samples. By leveraging the separability between the point process representation and the functional mapping, CliPS simultaneously accomplishes cluster identification, assessment of solution quality, and structural validation within the posterior space. The method effectively extracts distinguishable clustering patterns by isolating interpretable features from complex posterior geometries. Extensive experiments on both simulated and real-world datasets demonstrate that CliPS reliably recovers well-separated cluster distributions, confirming its effectiveness and broad applicability across diverse data scenarios.

Bayesian mixture modelscluster identificationcluster validation

Latest Papers

What's happening recently
View more

This study addresses the joint inference of a consensus ranking, individual preferences, and clustering structure from preference data without requiring a pre-specified number of clusters. To this end, it introduces Bayesian nonparametrics into the Mallows model for the first time, constructing an extended framework based on Dirichlet process mixtures that accommodates both incomplete rankings and pairwise comparison data. The proposed approach enables simultaneous posterior inference of the number of clusters and cluster assignments via Markov chain Monte Carlo (MCMC) sampling. Experimental results demonstrate that the method accurately recovers the true number of clusters in simulated data, outperforming finite mixture models, and significantly enhances personalized recommendation performance on real-world movie rating data, particularly excelling in predicting missing ratings.

Bayesian nonparametricsclusteringDirichlet process

This study addresses the need for precise identification of disease subtypes in clinical decision-making by proposing a Bayesian nonparametric clustering method based on the Dirichlet process mixture model. The approach leverages coordinate-ascent variational inference to enable efficient patient stratification, integrating variational inference into a nonparametric Bayesian framework to significantly reduce computational complexity while maintaining high clustering accuracy and mitigating misdiagnosis risks. Experimental results demonstrate that the model accurately recovers ground-truth cluster structures in synthetic data, achieving superior performance in homogeneity and completeness metrics, and exhibits substantially improved computational efficiency compared to conventional Markov chain Monte Carlo (MCMC) methods.

Bayesian nonparametricsDirichlet process mixture modeldisease subtypes

This work proposes a Bayesian nonparametric clustering method for replicated marked Poisson point process data, jointly inferring the latent cluster structure, the number of clusters, and the intensity surface associated with continuous marks. Built upon a Dirichlet process mixture model, the approach employs a squared link function to model the intensity surface and leverages variational Bayesian inference for efficient learning. To address sign ambiguity and nodal line issues inherent in the squared link, the method introduces a constrained Laplace approximation that reformulates the non-conjugate basis coefficient updates as a constrained optimization problem, thereby providing theoretical guarantees for mode-finding. Experimental results demonstrate that the proposed model achieves superior performance in clustering accuracy, intensity estimation, and computational efficiency on both synthetic and real-world datasets.

Bayesian nonparametricsclusteringDirichlet process mixture

This study presents the first systematic investigation into the statistical properties of the Dirichlet process when employed as a sampling distribution and introduces a Bayesian inference framework for its base measure and concentration parameter. Treating the Dirichlet process as a data-generating mechanism, the authors develop a joint inference approach for these two key parameters by integrating Bayesian nonparametric modeling with Markov chain Monte Carlo algorithms, leveraging observed histogram sequences. The proposed methodology is validated through extensive experiments on both synthetic and real-world datasets, demonstrating its effectiveness and practical utility. This work addresses a notable gap in the literature by providing a principled solution to parameter inference in Dirichlet process models, thereby advancing the theoretical and applied understanding of this foundational Bayesian nonparametric construct.

Bayesian inferencecentering measureDirichlet Process

Hot Scholars

DA

David A. Stephens

Professor, Department of Mathematics and Statistics, McGill University
Statistics
KN

Khai Nguyen

Ph.D. Candidate at University of Texas at Austin
Machine LearningOptimal TransportBayesian StatisticsBioinformatics
DP

Debdeep Pati

Professor, Department of Statistics, University of Wisconsin - Madison
Bayesian nonparametricshigh-dimensional data analysis
MW

Mengzhu Wang

National University of Defense Technology
transfer learningcomputer vision
YW

Yingxu Wang

Mohamed bin Zayed University of Artificial Intelligence
Graph LearningAI4Science