perform subgroup analysis

Designs and executes analyses to identify, discover, and interpret subgroups or clusters within a population (including clustering and patient subtyping), using stratification and clustering methods to define group membership. Builds estimators and sampling procedures to measure subgroup-specific metrics and heterogeneity, compare performance across groups, and test calibration, bias, and uncertainty (including posterior or resampling-based inference) for informed interpretation and downstream decision-making.

performsubgroupanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Subgroup Performance Analysis in Hidden Stratifications

Mar 13, 2025
AB
Alceu Bissoto
🏛️ UDEM | University of Bern | Diabetes Center Berne | University of Tübingen | University of Lucerne

Medical AI models often exhibit implicit performance disparities across real-world patient populations, yet conventional subgroup analyses—relying on limited, predefined metadata (e.g., sex)—fail to uncover their root causes. To address this, we propose the first unsupervised implicit subgroup discovery framework for trustworthy medical AI deployment: it requires neither labels nor prior metadata, instead identifying performance-sensitive patient subgroups via feature-space clustering and quantifying diagnostic performance gaps across them. Our methodological innovations include a lightweight, scalable subgroup discovery pipeline and a novel performance disparity assessment strategy grounded in calibrated uncertainty estimation. Evaluated on chest X-ray and skin lesion classification tasks, the framework reveals inter-subgroup accuracy gaps exceeding 30%—entirely undetected by standard metadata-based analysis. This provides the first empirically validated, interpretable framework enabling fine-grained model validation and continuous monitoring in clinical AI deployment.

Discover hidden stratifications using learned feature representations.Evaluate subgroup discovery for medical AI performance monitoring.Identify performance disparities in ML models across patient groups.

This study addresses the limitations of existing exhaustive subgroup treatment effect plots, which struggle to reliably assess heterogeneity under small sample sizes and multiple testing, and lack a formal quantification of the significance of observed heterogeneity under the null hypothesis of homogeneous treatment effects. The authors propose a computationally efficient strategy to construct homogeneity regions by leveraging a Doubly Robust learner to generate pseudo-outcomes for subgroup effect estimation. By constructing a reference distribution under homogeneity, the method provides the first framework to quantify evidence of heterogeneity directly within exhaustive subgroup plots. An explicit formula for homogeneity regions is derived, accompanied by several approaches for computing critical thresholds. Empirical evaluations in cardiovascular clinical trials and simulation studies demonstrate well-calibrated performance and substantial improvements over conventional methods based on subgroup mean differences.

clinical trialexploratory subgroup analysisheterogeneity

Identifying treatment response subgroups in observational time-to-event data

Aug 06, 2024
VJ
Vincent Jeanselme
🏛️ University of Cambridge | The Alan Turing Institute | University of Oxford

Identifying heterogeneous treatment effects (HTE) in observational time-to-event data remains challenging, and conventional randomized controlled trial (RCT) subgroup analysis methods often yield biased estimates in real-world settings. Method: We propose the first outcome-oriented, dynamic subgroup discovery framework for causal survival analysis. Our approach jointly models the covariate–treatment–outcome triad by integrating doubly robust causal inference, Cox-type survival modeling, and interpretable clustering—enabling both individualized treatment effect estimation and average treatment effect calibration. Contribution/Results: Evaluated on multi-center RCT and observational cohort datasets, our method significantly outperforms state-of-the-art baselines in identifying clinically meaningful responder subgroups with high precision. It bridges the evidence gap between RCTs and real-world practice, delivering interpretable, generalizable subgroup insights to support clinical guideline development and personalized decision-making.

Identifying treatment response subgroupsNovel strategy for RCT and observational studiesOvercoming RCT limitations in subgroup analysis

Data-driven controlled subgroup selection in clinical trials

Dec 17, 2025
MM
Manuel M. Müller
🏛️ University of Cambridge | Novartis Pharma AG | University of Edinburgh | University of Southampton | Nanjing University | Lancaster University

Data-driven subgroup identification in clinical trials suffers from post-selection inference issues, leading to inflated Type I error rates and biased effect estimates—hindering the implementation of precision medicine. To address the dual objective of identifying both *safe subgroups* (with low adverse event risk) and *efficacious subgroups* (with high treatment effect), this paper proposes two novel controlled subgroup selection methods: one based on generalized linear models and another within an isotonic regression framework. For the first time in a regression setting, both methods enable rigorous post-selection inference with guaranteed Type I error control under the null. Comprehensive simulation studies demonstrate robust error rate control across diverse scenarios and quantify sensitivity to modeling assumptions. The proposed methods provide a statistically rigorous, reproducible, and practically applicable toolkit for clinical subgroup analysis.

Addresses post-selection inference to control Type I error ratesDevelops methods for selecting patient subgroups in clinical trialsIdentifies subgroups with high treatment effect or safety from adverse events

Statistical Performance Guarantee for Subgroup Identification with Generic Machine Learning

Oct 12, 2023
ML
Michael Lingzhi Li
🏛️ Harvard Business School | Harvard University

In causal subgroup identification, conventional methods suffer from high estimation noise in conditional average treatment effect (CATE) estimation and multiplicity issues arising from two-stage procedures. To address these challenges, this paper proposes the Global Adaptive Treatment Effect Sets (GATES) uniform confidence band method. Grounded in randomized trial design and empirical process theory, GATES provides finite-sample, model-agnostic global statistical guarantees for CATE estimates produced by arbitrary black-box machine learning models—without requiring parametric assumptions or resampling. It enables rigorous, threshold-agnostic identification of credible subgroups exhibiting clinically meaningful treatment effects. Empirically, GATES maintains nominal coverage even in small samples (n = 100), substantially improving the reliability of subgroup inference. Applied to a late-stage prostate cancer clinical trial, it robustly identifies a clinically significant “exceptional responder” subgroup. This work establishes a verifiable, statistically principled paradigm for causal subgroup discovery in precision medicine.

Addressing bias and noise in CATE estimation for subgroup identificationAvoiding modeling assumptions and intensive resampling proceduresProviding statistical guarantees for treatment effect subgroup selection

Latest Papers

What's happening recently
View more

This study addresses the selection bias inherent in data-driven subgroup discovery within clinical trials, which inflates standard effect estimates and invalidates confidence intervals. To mitigate this issue, the authors propose an algorithm-agnostic post-selection inference framework that decouples subgroup identification from effect reporting. By employing refitting-free multiplier resampling via a joint perturbation technique, they establish the coverage properties of conditionally adaptive estimators. Integrating forest search with causal forests, the proposed approach substantially alleviates selection bias in both simulation studies and real-world trial applications. The resulting confidence intervals demonstrate superior coverage compared to conventional bootstrap methods, thereby providing rigorous statistical guarantees for subgroup analysis.

clinical trialsdata-driven subgroup discoverypost-selection inference

Traditional subgroup analyses in observational biomedical data often yield unstable and difficult-to-interpret results due to individuals experiencing only a single exposure, non-identifiable true causal effects, and uncertain confounding structures. This work proposes an integrated framework that first selects covariates via causal discovery, then constructs exposure- and outcome-agnostic pretreatment subgroups using unsupervised clustering methods—including K-means, fuzzy C-means, and Bayesian Gaussian mixture models. Subsequently, it evaluates hypothetical intervention strategies through uncertainty-aware screening combined with doubly robust estimation. The approach uniquely unifies unsupervised subgroup discovery with policy evaluation and introduces empirical Bernstein gating and Bayesian pooling to control risk. Applied to PIMA and NHANES datasets, the optimal policies achieved utilities of 0.735–0.799, though risk differences became nonsignificant after multiple testing correction.

causal inferenceobservational datapolicy prioritization

This study addresses the modeling challenges posed by population heterogeneity in high-dimensional clinical data by systematically reviewing and categorizing methods that integrate patient covariate clustering with outcome modeling. It explicitly distinguishes, for the first time, between “informed clustering” (which leverages outcome information) and “agnostic clustering” (based solely on covariates). Through a comprehensive analysis of 55 studies—including PPMx models, finite mixture regression, cluster-aware supervised learning, and two-stage approaches—the work clarifies the respective strengths and appropriate use cases of each framework in risk stratification, subgroup treatment effect estimation, and rare disease research. The review provides a clear methodological guide and practical reference for integrated modeling of heterogeneous clinical data.

clusteringheterogeneous populationshigh-dimensional covariates

This study addresses the failure of inference for data-driven subgroup identification in within-sample evaluation due to selection bias—particularly when subgroup boundaries are non-smooth and depend on infinite-dimensional functionals. The authors propose a conditional adaptive perturbation method grounded in a triple robustness theoretical framework, which accommodates any machine learning algorithm, including black-box models, without requiring parametric assumptions or smoothness conditions on subgroup boundaries. The approach jointly optimizes subgroup identification and nuisance parameter estimation rates, enabling fully efficient, unbiased within-sample inference without data splitting. In a reanalysis of the ACTG 175 clinical trial, the method substantially improves estimation stability and statistical efficiency while avoiding the information loss inherent in conventional sample-splitting strategies.

data-dependent objectsin-sample evaluationnonregularity

This study addresses the lack of rigorous statistical assessment for the reliability of output structures in complex clustering pipelines that involve multiple data-dependent stages such as anomaly detection, feature selection, and clustering. To bridge this gap, the work systematically applies selective inference to the entire clustering analysis workflow, establishing a statistical framework that enables valid significance testing of final cluster assignments. The proposed method rigorously controls the type I error rate at any pre-specified nominal level and demonstrates strong empirical performance on both synthetic and real-world datasets. By doing so, it provides a principled and reliable foundation for statistical inference in multi-stage, data-driven clustering procedures.

clustering pipelinesdata analysis pipelineselective inference

Hot Scholars

BG

Ben Glocker

Imperial College London
Medical Image AnalysisComputer VisionMachine Learning
RA

Ritu Agarwal

Johns Hopkins Carey Business School
Healthcare information technologyhealth analyticsAI in healthcareinnovation adoption and diffusion
HT

Hari Trivedi

Emory University
Deep LearningRadiologyMammographyAI
CB

Conrad Borchers

Carnegie Mellon University
Educational Data MiningLearning AnalyticsIntelligent Tutoring SystemsSelf-Regulated Learning