Score
Designs and executes analyses to identify, discover, and interpret subgroups or clusters within a population (including clustering and patient subtyping), using stratification and clustering methods to define group membership. Builds estimators and sampling procedures to measure subgroup-specific metrics and heterogeneity, compare performance across groups, and test calibration, bias, and uncertainty (including posterior or resampling-based inference) for informed interpretation and downstream decision-making.
This study addresses the heterogeneity of treatment effects in clinical research, where conventional subgroup analyses lack individual-level predictive power and purely machine learning–based approaches often lack statistical guarantees. To bridge this gap, the authors propose a two-stage hybrid workflow: first, using formal statistical hypothesis testing to confirm the presence of heterogeneous treatment effects, then constructing an individualized treatment strategy evaluated via cross-fitted doubly robust estimation under a Neyman–Pearson risk constraint. This framework integrates the interpretability of statistical inference with the predictive strength of machine learning, yielding a transparent, auditable, and statistically principled approach to heterogeneity. The method demonstrates efficacy in both simulation studies and the ACTG 175 HIV trial, and is accompanied by a practical implementation checklist along with guidance for alignment with regulatory-oriented heterogeneous treatment effect (HTE) assessment protocols.
Medical AI models often exhibit implicit performance disparities across real-world patient populations, yet conventional subgroup analyses—relying on limited, predefined metadata (e.g., sex)—fail to uncover their root causes. To address this, we propose the first unsupervised implicit subgroup discovery framework for trustworthy medical AI deployment: it requires neither labels nor prior metadata, instead identifying performance-sensitive patient subgroups via feature-space clustering and quantifying diagnostic performance gaps across them. Our methodological innovations include a lightweight, scalable subgroup discovery pipeline and a novel performance disparity assessment strategy grounded in calibrated uncertainty estimation. Evaluated on chest X-ray and skin lesion classification tasks, the framework reveals inter-subgroup accuracy gaps exceeding 30%—entirely undetected by standard metadata-based analysis. This provides the first empirically validated, interpretable framework enabling fine-grained model validation and continuous monitoring in clinical AI deployment.
This study addresses the limitations of existing exhaustive subgroup treatment effect plots, which struggle to reliably assess heterogeneity under small sample sizes and multiple testing, and lack a formal quantification of the significance of observed heterogeneity under the null hypothesis of homogeneous treatment effects. The authors propose a computationally efficient strategy to construct homogeneity regions by leveraging a Doubly Robust learner to generate pseudo-outcomes for subgroup effect estimation. By constructing a reference distribution under homogeneity, the method provides the first framework to quantify evidence of heterogeneity directly within exhaustive subgroup plots. An explicit formula for homogeneity regions is derived, accompanied by several approaches for computing critical thresholds. Empirical evaluations in cardiovascular clinical trials and simulation studies demonstrate well-calibrated performance and substantial improvements over conventional methods based on subgroup mean differences.
Identifying heterogeneous treatment effects (HTE) in observational time-to-event data remains challenging, and conventional randomized controlled trial (RCT) subgroup analysis methods often yield biased estimates in real-world settings. Method: We propose the first outcome-oriented, dynamic subgroup discovery framework for causal survival analysis. Our approach jointly models the covariate–treatment–outcome triad by integrating doubly robust causal inference, Cox-type survival modeling, and interpretable clustering—enabling both individualized treatment effect estimation and average treatment effect calibration. Contribution/Results: Evaluated on multi-center RCT and observational cohort datasets, our method significantly outperforms state-of-the-art baselines in identifying clinically meaningful responder subgroups with high precision. It bridges the evidence gap between RCTs and real-world practice, delivering interpretable, generalizable subgroup insights to support clinical guideline development and personalized decision-making.
Data-driven subgroup identification in clinical trials suffers from post-selection inference issues, leading to inflated Type I error rates and biased effect estimates—hindering the implementation of precision medicine. To address the dual objective of identifying both *safe subgroups* (with low adverse event risk) and *efficacious subgroups* (with high treatment effect), this paper proposes two novel controlled subgroup selection methods: one based on generalized linear models and another within an isotonic regression framework. For the first time in a regression setting, both methods enable rigorous post-selection inference with guaranteed Type I error control under the null. Comprehensive simulation studies demonstrate robust error rate control across diverse scenarios and quantify sensitivity to modeling assumptions. The proposed methods provide a statistically rigorous, reproducible, and practically applicable toolkit for clinical subgroup analysis.
In causal subgroup identification, conventional methods suffer from high estimation noise in conditional average treatment effect (CATE) estimation and multiplicity issues arising from two-stage procedures. To address these challenges, this paper proposes the Global Adaptive Treatment Effect Sets (GATES) uniform confidence band method. Grounded in randomized trial design and empirical process theory, GATES provides finite-sample, model-agnostic global statistical guarantees for CATE estimates produced by arbitrary black-box machine learning models—without requiring parametric assumptions or resampling. It enables rigorous, threshold-agnostic identification of credible subgroups exhibiting clinically meaningful treatment effects. Empirically, GATES maintains nominal coverage even in small samples (n = 100), substantially improving the reliability of subgroup inference. Applied to a late-stage prostate cancer clinical trial, it robustly identifies a clinically significant “exceptional responder” subgroup. This work establishes a verifiable, statistically principled paradigm for causal subgroup discovery in precision medicine.
This study addresses the selection bias inherent in data-driven subgroup discovery within clinical trials, which inflates standard effect estimates and invalidates confidence intervals. To mitigate this issue, the authors propose an algorithm-agnostic post-selection inference framework that decouples subgroup identification from effect reporting. By employing refitting-free multiplier resampling via a joint perturbation technique, they establish the coverage properties of conditionally adaptive estimators. Integrating forest search with causal forests, the proposed approach substantially alleviates selection bias in both simulation studies and real-world trial applications. The resulting confidence intervals demonstrate superior coverage compared to conventional bootstrap methods, thereby providing rigorous statistical guarantees for subgroup analysis.
Traditional subgroup analyses in observational biomedical data often yield unstable and difficult-to-interpret results due to individuals experiencing only a single exposure, non-identifiable true causal effects, and uncertain confounding structures. This work proposes an integrated framework that first selects covariates via causal discovery, then constructs exposure- and outcome-agnostic pretreatment subgroups using unsupervised clustering methods—including K-means, fuzzy C-means, and Bayesian Gaussian mixture models. Subsequently, it evaluates hypothetical intervention strategies through uncertainty-aware screening combined with doubly robust estimation. The approach uniquely unifies unsupervised subgroup discovery with policy evaluation and introduces empirical Bernstein gating and Bayesian pooling to control risk. Applied to PIMA and NHANES datasets, the optimal policies achieved utilities of 0.735–0.799, though risk differences became nonsignificant after multiple testing correction.
This study addresses the modeling challenges posed by population heterogeneity in high-dimensional clinical data by systematically reviewing and categorizing methods that integrate patient covariate clustering with outcome modeling. It explicitly distinguishes, for the first time, between “informed clustering” (which leverages outcome information) and “agnostic clustering” (based solely on covariates). Through a comprehensive analysis of 55 studies—including PPMx models, finite mixture regression, cluster-aware supervised learning, and two-stage approaches—the work clarifies the respective strengths and appropriate use cases of each framework in risk stratification, subgroup treatment effect estimation, and rare disease research. The review provides a clear methodological guide and practical reference for integrated modeling of heterogeneous clinical data.
This study addresses the failure of inference for data-driven subgroup identification in within-sample evaluation due to selection bias—particularly when subgroup boundaries are non-smooth and depend on infinite-dimensional functionals. The authors propose a conditional adaptive perturbation method grounded in a triple robustness theoretical framework, which accommodates any machine learning algorithm, including black-box models, without requiring parametric assumptions or smoothness conditions on subgroup boundaries. The approach jointly optimizes subgroup identification and nuisance parameter estimation rates, enabling fully efficient, unbiased within-sample inference without data splitting. In a reanalysis of the ACTG 175 clinical trial, the method substantially improves estimation stability and statistical efficiency while avoiding the information loss inherent in conventional sample-splitting strategies.
This study addresses the lack of rigorous statistical assessment for the reliability of output structures in complex clustering pipelines that involve multiple data-dependent stages such as anomaly detection, feature selection, and clustering. To bridge this gap, the work systematically applies selective inference to the entire clustering analysis workflow, establishing a statistical framework that enables valid significance testing of final cluster assignments. The proposed method rigorously controls the type I error rate at any pre-specified nominal level and demonstrates strong empirical performance on both synthetic and real-world datasets. By doing so, it provides a principled and reliable foundation for statistical inference in multi-stage, data-driven clustering procedures.