Score
Designs and fits Bayesian latent-class or finite-mixture models that represent a population as a mixture of discrete latent groups, estimating class‑specific parameter distributions and posterior probabilities of class membership from observed data (categorical, continuous, or mixed). Performs model selection (e.g., number of classes), quantifies uncertainty in class assignments and parameters via the posterior, and implements extensions such as covariate-dependent classes, hierarchical structure, and posterior predictive checks.
This study addresses three core challenges in Bayesian random partition models: modeling uncertainty in the number of clusters, constructing appropriate priors, and summarizing posterior distributions. To overcome theoretical inconsistencies in existing cluster-number estimation methods, we systematically classify and evaluate the computational tractability and statistical validity of partition priors—including the Pitman–Yor process and Chinese Restaurant Process—for the first time. We propose a novel paradigm for posterior inference that combines dimensionality reduction in the partition space with structured posterior summarization, integrating MCMC sampling, variational inference, and posterior consistency analysis. Furthermore, we establish the first comprehensive methodological evaluation framework, identifying critical theoretical bottlenecks. Our work provides an interpretable, scalable, and adaptive clustering methodology tailored for high-dimensional data applications—such as genomics and medical imaging—while ensuring principled uncertainty quantification and model selection.
Gibbs sampling for Bayesian mixture models suffers from slow mixing in the marginal posterior over component assignments and struggles to jointly perform model selection and parameter inference. Method: We propose two novel joint-sampling MCMC algorithms: (1) a collapsed Gibbs sampler incorporating unconventional move sets, and (2) a prior-driven, rejection-free component allocation sampler. Both methods jointly update observation assignments and the number of components, unifying model fitting and dimensionality inference. Contribution/Results: Our approaches eliminate the need for post-hoc model selection and substantially improve Markov chain mixing efficiency. In latent class analysis tasks, they reduce mixing time by several-fold compared to state-of-the-art methods while achieving comparable or superior posterior inference accuracy. The framework provides an efficient, fully automated computational solution for high-dimensional Bayesian nonparametric modeling.
Multivariate longitudinal data—common in physiological and financial health research—frequently exhibit zero inflation, posing challenges for conventional modeling. To address this, we propose a Bayesian latent class mixture model that unifies Tobit, two-part, and zero-inflated Poisson (ZIP) structures, enabling flexible characterization of distinct zero-generating mechanisms. Innovatively, we couple an adaptive Lasso-type shrinkage prior to simultaneously perform variable selection and latent class identification, thereby capturing complex dependencies among multivariate trajectories and heterogeneous population substructures. Applied to the Health and Retirement Study (HRS) data, the model successfully identifies clinically meaningful subtypes exhibiting coupled physiological–financial decline. Simulation studies demonstrate that our approach significantly outperforms existing methods in both parameter estimation accuracy and latent class assignment precision.
This paper addresses two fundamental challenges in latent class modeling for high-dimensional binary data: identifying individual class memberships and automatically determining the number of latent classes. We propose a two-stage algorithm comprising spectral clustering initialization followed by a single-step maximum likelihood refinement. Theoretically, under mild regularity conditions, the method achieves optimal latent class recovery and exact clustering consistency. Moreover, we construct a simple, consistent, and tuning-free estimator for the number of latent classes. Extensive simulations and real-data analyses demonstrate that our approach significantly outperforms existing methods in recovery accuracy, computational efficiency, and statistical consistency. Crucially, it offers both rigorous theoretical guarantees—establishing optimality and consistency—and strong practical utility, making it well-suited for high-dimensional binary data analysis.
This work addresses the challenge that irrelevant variables in high-dimensional mixture models can obscure the identification of latent subgroups. To overcome this, the authors propose a Bayesian distillation clustering framework that tightly integrates variable selection with mixture model structure. By leveraging Bayesian variable selection and posterior inclusion probabilities, the method identifies a discriminative subset of covariates crucial for subgroup separation and performs clustering within the resulting low-dimensional, statistically interpretable subspace. The approach further incorporates control of the expected false discovery proportion to ensure scientific validity and employs conditional independence diagnostics to uncover the dependence structure among variables within each subgroup. This integrated strategy substantially enhances the accuracy, interpretability, and scientific relevance of identifying heterogeneous subgroups in high-dimensional settings.
This paper addresses longitudinal polytomous response data featuring ordinal attributes and individual-level covariates. We propose a Restricted Latent-Class Hidden Markov Model (RLC-HMM) that jointly models the evolution of latent attributes and the response-generating process, accommodating time non-homogeneity and conditional dependencies among states and covariates. To our knowledge, this is the first rigorous proof of model identifiability under such a complex structure. By integrating latent-class analysis with covariate-dependent state transitions, the RLC-HMM supports exploratory longitudinal cognitive diagnosis. Bayesian inference via MCMC is employed; simulations confirm accurate and robust parameter estimation. An empirical application to mathematics assessment data demonstrates substantial improvements over existing confirmatory approaches, effectively uncovering dynamic developmental trajectories of student proficiency.
This study addresses the limitations of covariate logistic models in latent class analysis, specifically their difficulty in capturing complex interactions and providing sufficient interpretability. To overcome these challenges, this work proposes a novel framework that directly integrates decision trees into latent class analysis. By leveraging interpretable tree structures combined with pruning and binary splitting techniques, the method effectively models covariate effects without requiring additional assumptions. Empirical validation demonstrates the approach's efficacy, yielding highly interpretable classification paths. Consequently, this research significantly expands the methodological toolkit for latent class analysis, establishing a new paradigm for handling complex covariate relationships while enhancing model transparency and practical utility.
This work addresses the challenges of parameter explosion and difficulty in modeling inter-class joint dependencies in high-dimensional categorical space classification. The authors propose an identifiable reduced-rank spatial multinomial model that captures class-specific spatial effects through shared low-dimensional latent factors, substantially reducing parameter dimensionality while preserving the dependence structure among categories. To overcome the failure of conventional conjugate priors and Pólya-Gamma augmentation under this factorized formulation, they develop a Gibbs sampling algorithm incorporating Metropolis–Hastings updates based on Laplace approximations. Simulations demonstrate the effectiveness of the proposed dimension selection and proposal distributions, and the method enables scalable inference in mapping dominant tree species across the Blue Ridge Mountains, supporting flexible predictions for individual classes, unions of classes, and aggregated regional summaries.
This study addresses the challenge of latent class modeling for mixed continuous and binary data by proposing a unified joint-likelihood framework that models continuous variables with normal distributions and binary variables with Bernoulli distributions. The latent class structure is efficiently estimated via the EM algorithm. A key contribution is the development of the first frequentist R package tailored to such mixed-data latent class models, which accommodates heteroscedasticity and censoring, eliminates the need for manual likelihood derivation by users, and includes dedicated tools for summarization and visualization. The method has been successfully applied to EQ-5D-5L value set estimation, demonstrating its practical utility and ease of use in health economics and related fields.
This study addresses the challenge of inferring underlying population distributions from aggregated data available only in the form of histograms or frequency tables. The authors propose a nonparametric Bayesian inference method based on mixture models, employing reversible-jump Markov chain Monte Carlo to fit Gaussian mixtures with either a finite or countably infinite number of components. This work represents the first systematic application of a nonparametric Bayesian framework to histogram data analysis. Furthermore, by leveraging Dirichlet processes, the approach jointly models multiple histograms, enabling information sharing across groups and providing posterior probabilities to quantify homogeneity among them. Empirical evaluations demonstrate that the method effectively reconstructs complex distributions from large-scale aggregated data and offers principled clustering and homogeneity assessment.
Latent class analysis (LCA) often suffers from limited interpretability due to the complexity of item response probability matrices. This work proposes a post-estimation sparsification method that enhances traditional LCA by incorporating item-wise pseudolikelihood with sparse regularization, which penalizes the number of non-zero response probability levels per item. This approach automatically merges redundant response levels, thereby improving the interpretability of latent classes. The method is computationally efficient and enjoys theoretical consistency guarantees for accurately recovering the true sparse structure. Combined with Bayesian Information Criterion (BIC) for selecting the number of latent classes, the proposed framework demonstrates strong performance in both simulation studies and an empirical analysis of social role performance survey data, yielding concise and interpretable latent class characterizations. The implementation code is publicly available.