π€ AI Summary
Bayesian clustering quantifies posterior uncertainty but lacks a general method to summarize uncertainty over the partition space. This paper proposes a generic post-processing framework for Markov chain Monte Carlo (MCMC) posterior samples: it constructs the first interpretable and computationally efficient posterior credible set of clusterings; introduces a novel uncertainty measure enabling cumulative posterior estimation and credible region construction for cluster-specific parameters; and operates entirely without point estimates of partitions or label alignment. By integrating statistical modeling on the partition space with optimization-based aggregation, the method enhances interpretability, robustness, and practical utility across multiple empirical studies. It establishes a reliable, off-the-shelf paradigm for uncertainty quantification in Bayesian clustering.
π Abstract
Bayesian clustering methods have the widely touted advantage of providing a probabilistic characterization of uncertainty in clustering through the posterior distribution. An amazing variety of priors and likelihoods have been proposed for clustering in a broad array of settings. There is also a rich literature on Markov chain Monte Carlo (MCMC) algorithms for sampling from posterior clustering distributions. However, there is relatively little work on summarizing the posterior uncertainty. The complexity of the partition space corresponding to different clusterings makes this problem challenging. We propose a post-processing procedure for any Bayesian clustering model with posterior samples that generates a credible set that is easy to use, fast to compute, and intuitive to interpret. We also provide new measures of clustering uncertainty and show how to compute cluster-specific parameter estimates and credible regions that accumulate a desired posterior probability without having to condition on a partition estimate or employ label-switching techniques. We illustrate our procedure through several applications.