Inference for Standard Trial Estimands after Data-Driven Subgroup Discovery

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the selection bias inherent in data-driven subgroup discovery within clinical trials, which inflates standard effect estimates and invalidates confidence intervals. To mitigate this issue, the authors propose an algorithm-agnostic post-selection inference framework that decouples subgroup identification from effect reporting. By employing refitting-free multiplier resampling via a joint perturbation technique, they establish the coverage properties of conditionally adaptive estimators. Integrating forest search with causal forests, the proposed approach substantially alleviates selection bias in both simulation studies and real-world trial applications. The resulting confidence intervals demonstrate superior coverage compared to conventional bootstrap methods, thereby providing rigorous statistical guarantees for subgroup analysis.
📝 Abstract
Data-driven subgroup identification is increasingly used in clinical trials, yet the effect reported for a discovered subgroup is typically the standard analysis: a Cox or generalized-linear-model treatment coefficient fitted within that subgroup. Selection by an effect-driven rule --- rewarding large estimated effects subject to screening and size criteria --- inflates the reported coefficient: a winner's curse invalidating naive intervals whether or not the identifier's internal estimand was de-biased by cross-fitting. We develop a procedure-agnostic framework for post-selection inference that separates discovery from reporting: forest search, difference-in-natural-parameters estimators, and causal forests generate interpretable candidate subgroups; selection is aligned with the coefficient to be reported; and the target is the standard-analysis effect of the selected subgroup, the candidate family held fixed --- a conditional, data-adaptive estimand. Resampling the search then reduces to jointly perturbing all candidate effects, yielding refit-free multiplier resampling with an infinitesimal-jackknife interval whose coverage we characterize under standard regularity and an explicit condition on the limiting competition. It matches the full bootstrap on a fixed family; on model-generated families it remains valid for the same conditional target, which the bootstrap reproduces only under further conditions. Two trial applications and matched simulations show mitigated selection bias and improved coverage.
Problem

Research questions and friction points this paper is trying to address.

post-selection inference
data-driven subgroup discovery
winner's curse
clinical trials
selection bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

post-selection inference
winner's curse
multiplier resampling
infinitesimal jackknife
data-driven subgroup discovery
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Larry F. Leon
K
Keaven M. Anderson