๐ค AI Summary
This study addresses the trade-off between computational efficiency and full posterior inference in Bayesian clustering of multivariate binary data by proposing a Bayesian mixture model that integrates a penalized complexity prior with an asymmetric Dirichlet prior. The approach accommodates a large number of latent components while enabling intuitive control over the distribution of the number of clusters through its asymmetric prior structure. Computational feasibility is ensured via an efficient Markov chain Monte Carlo (MCMC) algorithm. Empirical evaluations on both simulated and real-world ecological presenceโabsence species data demonstrate that the proposed model performs comparably or superiorly to existing methods, successfully achieving a balance among computational efficiency, Bayesian inferential completeness, and interpretability in cluster analysis.
๐ Abstract
Clustering multivariate binary data is of interest in many scientific fields, including ecology, biomedicine, and social policy. Beyond heuristic clustering algorithms, such data can be modelled using multivariate Bernoulli mixture models. Many Bayesian implementations of these models involve a trade-off between computational efficiency and full posterior inference. We propose instead a Bayesian approach able to provide both aspects. The method fixes the total number of components to a large value and employs an asymmetric Dirichlet prior on the mixture weights. The asymmetric Dirichlet hyperparameters are elicited using the popular Penalized Complexity prior framework, which provides an intuitive way for users to inform the induced distribution of the number of clusters. An efficient MCMC algorithm is then developed to fit the model. Simulations and real-world applications demonstrate that the method is competitive with existing alternatives and can outperform them in certain settings. The proposal is illustrated using an ecological dataset about presence-absence of species across multiple sites, where cluster-specific parameters are modelled on the basis of environmental conditions. Overall, the proposed method provides a computationally efficient, fully Bayesian, and interpretable framework for clustering multivariate binary data, with potential applications across diverse scientific domains.