🤖 AI Summary
This work addresses the challenge that irrelevant variables in high-dimensional mixture models can obscure the identification of latent subgroups. To overcome this, the authors propose a Bayesian distillation clustering framework that tightly integrates variable selection with mixture model structure. By leveraging Bayesian variable selection and posterior inclusion probabilities, the method identifies a discriminative subset of covariates crucial for subgroup separation and performs clustering within the resulting low-dimensional, statistically interpretable subspace. The approach further incorporates control of the expected false discovery proportion to ensure scientific validity and employs conditional independence diagnostics to uncover the dependence structure among variables within each subgroup. This integrated strategy substantially enhances the accuracy, interpretability, and scientific relevance of identifying heterogeneous subgroups in high-dimensional settings.
📝 Abstract
Latent subgroup analysis is central to fields such as genomics, precision medicine, and social science, where the goal is to identify heterogeneous populations with distinct covariate structures or response behaviors. Mixture models provide a natural probabilistic framework for this task, representing the data-generating distribution as a weighted combination of subgroup-specific laws with unobserved labels. In high-dimensional regimes, these analyses face significant challenges. Often, only a small subset of covariates drives meaningful subgroup separation; the remaining variables may introduce noise or redundancy. Standard clustering methods typically treat all dimensions as equal, but in high-dimensional spaces, irrelevant coordinates can distort distances and obscure the low-dimensional structures defining latent classes. This paper introduces a Bayesian distilled clustering framework for high-dimensional mixture models. We propose that clustering should occur within a statistically justified subspace rather than the full ambient space. Our method utilizes a Bayesian variable selection model to estimate posterior inclusion probabilities, quantifying the evidence that each covariate contributes to subgroup separation or response behavior. A "distilled" covariate set is then identified by controlling the expected false-discovery proportion. Clustering is performed on this reduced subspace, followed by conditional independence diagnostics to examine subgroup-specific dependencies among selected variables. Critically, this framework is model-based: the distillation step is tied directly to the mixture structure and response model rather than a generic dimension-reduction criterion. This ensures the resulting subspace remains aligned with the scientific objective: identifying latent subgroups that differ in both distributional structure and behavior.