🤖 AI Summary
This work addresses the pervasive winner’s curse in data-driven selection, wherein post-selection effect sizes—such as those of optimal treatments, models, or features—are systematically overestimated. The authors propose a novel method that integrates adaptive exponential randomization with conditional inference to correct selection-induced bias in a nonparametric setting. By embedding an adaptive randomization mechanism within the conditional inference framework, the approach achieves substantially shorter confidence intervals while preserving selection quality comparable to standard top-k procedures. Demonstrated across diverse applications—including clinical treatment evaluation, model leaderboard inference, and feature importance analysis—the method effectively mitigates the winner’s curse, yielding tighter and more reliable statistical inference.
📝 Abstract
Researchers often select top-performing options or winners, based on a data-driven criterion, such as treatments, models, or model features and then report effect estimates for the selected winners. Naive post-selection estimates, however, are known to suffer from the winner's curse, producing systematically overoptimistic results. We introduce a flexible conditional inference method that corrects for this overoptimism through an adaptive exponential randomization scheme. Our method achieves selection quality that closely matches that of standard top-k selection, while also yielding shorter confidence intervals than existing approaches. Furthermore, our approach applies broadly to nonparametric settings with asymptotically linear selection statistics, covering wide-ranging applications such as inference for the efficacy of the most promising treatments in clinical trials, the abilities of top-ranked models on leaderboards, and the importance of the most predictive features in a model.