Score
Designs and fits finite mixture / latent class models that infer discrete unobserved subgroups and estimate class membership probabilities and class-specific parameters from multivariate observed data; models accommodate mixed continuous and categorical outcomes and are estimated with algorithms such as expectation–maximization and related likelihood-based methods.
This paper addresses two fundamental challenges in latent class modeling for high-dimensional binary data: identifying individual class memberships and automatically determining the number of latent classes. We propose a two-stage algorithm comprising spectral clustering initialization followed by a single-step maximum likelihood refinement. Theoretically, under mild regularity conditions, the method achieves optimal latent class recovery and exact clustering consistency. Moreover, we construct a simple, consistent, and tuning-free estimator for the number of latent classes. Extensive simulations and real-data analyses demonstrate that our approach significantly outperforms existing methods in recovery accuracy, computational efficiency, and statistical consistency. Crucially, it offers both rigorous theoretical guarantees—establishing optimality and consistency—and strong practical utility, making it well-suited for high-dimensional binary data analysis.
Multivariate longitudinal data—common in physiological and financial health research—frequently exhibit zero inflation, posing challenges for conventional modeling. To address this, we propose a Bayesian latent class mixture model that unifies Tobit, two-part, and zero-inflated Poisson (ZIP) structures, enabling flexible characterization of distinct zero-generating mechanisms. Innovatively, we couple an adaptive Lasso-type shrinkage prior to simultaneously perform variable selection and latent class identification, thereby capturing complex dependencies among multivariate trajectories and heterogeneous population substructures. Applied to the Health and Retirement Study (HRS) data, the model successfully identifies clinically meaningful subtypes exhibiting coupled physiological–financial decline. Simulation studies demonstrate that our approach significantly outperforms existing methods in both parameter estimation accuracy and latent class assignment precision.
This study addresses the challenge of latent class modeling for mixed continuous and binary data by proposing a unified joint-likelihood framework that models continuous variables with normal distributions and binary variables with Bernoulli distributions. The latent class structure is efficiently estimated via the EM algorithm. A key contribution is the development of the first frequentist R package tailored to such mixed-data latent class models, which accommodates heteroscedasticity and censoring, eliminates the need for manual likelihood derivation by users, and includes dedicated tools for summarization and visualization. The method has been successfully applied to EQ-5D-5L value set estimation, demonstrating its practical utility and ease of use in health economics and related fields.
This study addresses the substantial parameter bias incurred by maximum likelihood estimation (MLE) in finite mixture models under finite samples, particularly when component densities exhibit high overlap or possess unbounded support. To mitigate this issue, the authors propose a classification–mixture likelihood function grounded in a consistent classifier, yielding a parameter estimator with reduced bias. The method enhances finite-sample performance under relatively weak assumptions and achieves oracle efficiency under specific conditions. Theoretical analysis, corroborated by Monte Carlo simulations, demonstrates that the proposed estimator consistently outperforms standard MLE in both bias and mean squared error. Empirical application to panel latent class modeling of health administrative data shows a 17.6% reduction in out-of-sample prediction error compared to conventional MLE.
This paper addresses longitudinal polytomous response data featuring ordinal attributes and individual-level covariates. We propose a Restricted Latent-Class Hidden Markov Model (RLC-HMM) that jointly models the evolution of latent attributes and the response-generating process, accommodating time non-homogeneity and conditional dependencies among states and covariates. To our knowledge, this is the first rigorous proof of model identifiability under such a complex structure. By integrating latent-class analysis with covariate-dependent state transitions, the RLC-HMM supports exploratory longitudinal cognitive diagnosis. Bayesian inference via MCMC is employed; simulations confirm accurate and robust parameter estimation. An empirical application to mathematics assessment data demonstrates substantial improvements over existing confirmatory approaches, effectively uncovering dynamic developmental trajectories of student proficiency.
This study addresses the limitations of covariate logistic models in latent class analysis, specifically their difficulty in capturing complex interactions and providing sufficient interpretability. To overcome these challenges, this work proposes a novel framework that directly integrates decision trees into latent class analysis. By leveraging interpretable tree structures combined with pruning and binary splitting techniques, the method effectively models covariate effects without requiring additional assumptions. Empirical validation demonstrates the approach's efficacy, yielding highly interpretable classification paths. Consequently, this research significantly expands the methodological toolkit for latent class analysis, establishing a new paradigm for handling complex covariate relationships while enhancing model transparency and practical utility.
This work addresses the challenge of inefficient posterior exploration in hierarchical discrete models with latent variables, where conventional MCMC methods struggle due to the need to integrate out latent variables. The authors propose a similarity-driven MCMC approach that constructs a proposal mechanism based on a data-driven measure of discrepancy between observations and model predictions, thereby guiding transitions toward regions of higher posterior support without explicitly integrating latent variables. This method represents the first application of similarity-driven proposals to discrete-space MCMC and is naturally suited to complex hierarchical discrete models. Experiments on both synthetic and real-world data demonstrate substantial improvements in sampling efficiency and posterior exploration, confirming its effectiveness in models such as Dirichlet–Multinomial regression.
本文提出了一种非参数贝叶斯推理框架,用于解决部分识别的离散响应模型问题,通过直接对条件概率质量函数进行推理,避免了将条件矩转换为无条件矩或离散化协变量的需求。
本文针对未观察到的混杂因素问题,提出了一种混合学习视角来估计干预分布和因果效应,通过识别混合分布及成分机制实现。
This work addresses the challenge that irrelevant variables in high-dimensional mixture models can obscure the identification of latent subgroups. To overcome this, the authors propose a Bayesian distillation clustering framework that tightly integrates variable selection with mixture model structure. By leveraging Bayesian variable selection and posterior inclusion probabilities, the method identifies a discriminative subset of covariates crucial for subgroup separation and performs clustering within the resulting low-dimensional, statistically interpretable subspace. The approach further incorporates control of the expected false discovery proportion to ensure scientific validity and employs conditional independence diagnostics to uncover the dependence structure among variables within each subgroup. This integrated strategy substantially enhances the accuracy, interpretability, and scientific relevance of identifying heterogeneous subgroups in high-dimensional settings.