🤖 AI Summary
This study addresses the modeling challenges posed by population heterogeneity in high-dimensional clinical data by systematically reviewing and categorizing methods that integrate patient covariate clustering with outcome modeling. It explicitly distinguishes, for the first time, between “informed clustering” (which leverages outcome information) and “agnostic clustering” (based solely on covariates). Through a comprehensive analysis of 55 studies—including PPMx models, finite mixture regression, cluster-aware supervised learning, and two-stage approaches—the work clarifies the respective strengths and appropriate use cases of each framework in risk stratification, subgroup treatment effect estimation, and rare disease research. The review provides a clear methodological guide and practical reference for integrated modeling of heterogeneous clinical data.
📝 Abstract
This review provides a systematic overview of methods that combine covariate-based clustering of observational units (patients) with outcome models for clinical studies. We distinguish between informed-cluster models, where the outcome contributes to cluster formation, and agnostic-cluster models, where clustering is performed solely on covariates in a separate first step. Informed-cluster models include product partition models with covariates (PPMx), finite mixtures of regression models (FMR), and cluster-aware supervised learning (CluSL). Agnostic-cluster models encompass two-step procedures using either model-based or algorithmic clustering followed by cluster-specific regression models. Following a systematic search of Web of Science and PubMed, 55 records were identified that propose or evaluate such models. We describe the key models, summarise study characteristics, and present applications from biomedical and public health research. Clustering-based outcome models are particularly relevant for settings with high-dimensional covariates (e.g., biomarker panels and"omics") and heterogeneous patient populations. These models can support risk stratification and we discuss extensions to estimate subgroup-specific treatment effects. They are most valuable when the population is clustered in distinct regions of the covariate space that correspond to different outcome distributions. We discuss applications to rare disease research, covariate adjustment and borrowing from historical data, and subgroup-specific treatment effect estimation in clinical trials.