🤖 AI Summary
This study addresses the challenge that data-driven variable selection methods—such as Lasso and its adaptive variants—undermine the validity of classical statistical inference in Cox proportional hazards models, particularly leading to inflated false positive rates in right-censored survival data. The authors systematically evaluate the post-selection inference performance of sample splitting, exact post-selection inference, and debiased Lasso approaches. For the first time, these methods are comprehensively compared within a simulation framework designed to closely mimic real-world biomedical scenarios, and their practical utility is further validated using publicly available datasets. The findings elucidate the trade-offs among these methods in controlling Type I error rates and achieving estimation accuracy, thereby offering reliable and practical inference strategies for high-dimensional survival data analysis.
📝 Abstract
Choosing relevant predictors is central to the analysis of biomedical time-to-event data. Classical frequentist inference, however, presumes that the set of covariates is fixed in advance and does not account for data-driven variable selection. As a consequence, naive post-selection inference may be biased and misleading. In right-censored survival settings, these issues may be further exacerbated by the additional uncertainty induced by censoring. We investigate several inference procedures applied after variable selection for the coefficients of the Lasso and its extension, the adaptive Lasso, in the context of the Cox model. The methods considered include sample splitting, exact post-selection inference, and the debiased Lasso. Their performance is examined in a neutral simulation study reflecting realistic covariate structures and censoring rates commonly encountered in biomedical applications. To complement the simulation results, we illustrate the practical behavior of these procedures in an applied example using a publicly available survival dataset.