selection-aware inference

Designs and analyzes estimation and uncertainty-quantification procedures that remain valid after data-driven selection of models, features, or hyperparameters, producing selection-adjusted confidence intervals, p-values, and tuning rules that control selection-induced bias and error probabilities. This includes building adaptive hyperparameter-selection algorithms and calibration tests (e.g., learn-then-test, multiple-testing corrections, aggregation-of-candidate-tests, jackknife or other resampling calibrations) that guarantee conservative or finite-sample coverage and meet user-specified reliability constraints.

selection-awareinference

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Valid post-selection inference in Robust Q-learning

Aug 05, 2022
JJ
Jeremiah Jones
🏛️ Eli Lilly and Company | University of Pennsylvania | University of Rochester

In high-dimensional personalized treatment strategy estimation, standard post-variable-selection statistical inference fails due to selection-induced bias. Method: This paper introduces Universal Post-Selection Inference (UPoSI) into the robust Q-learning framework, proposing a selection-mechanism-agnostic universal post-selection inference method. The approach uniformly improves confidence interval construction and is theoretically shown to be asymptotically valid in multi-stage decision settings, guaranteeing nominal Type-I error control for hypothesis testing and exact coverage probability for confidence intervals. Results: Monte Carlo simulations demonstrate that the proposed method substantially improves coverage accuracy and statistical power compared to selective inference, while remaining compatible with diverse data-driven variable selection procedures. It thus provides a generalizable and verifiable foundation for statistical inference in robust Q-learning.

Addresses variable selection in robust Q-learning for treatment strategiesEnsures valid statistical inference after data-driven variable selectionHandles confounding and extraneous variables in adaptive treatment optimization

Adaptive Learn-then-Test: Statistically Valid and Efficient Hyperparameter Selection

Sep 24, 2024
MZ
Matteo Zecchin
🏛️ King's College London

In high-cost or high-risk settings where hyperparameter selection is constrained by limited test budgets and demanding statistical reliability requirements, this paper proposes the adaptive Learn-then-Test (aLTT) framework. Methodologically, aLTT introduces e-process theory—novelly applied to sequential, data-dependent multiple hypothesis testing—to enable rigorous false discovery rate (FDR) control with provably valid early stopping. Evaluated on offline reinforcement learning policy selection and large language model prompt optimization, aLTT achieves comparable statistical guarantees and final performance to classical LTT, while reducing the number of required test rounds by an order of magnitude. By unifying finite-sample statistical rigor with practical engineering efficiency, aLTT establishes a provably reliable, sample-efficient paradigm for AI model evaluation under resource constraints.

Hyperparameter OptimizationLimited Data SamplesStatistical Accuracy

Post-selection inference for quantifying uncertainty in changes in variance

May 24, 2024
RC
Rachel Carrington
🏛️ Lancaster University

Classical variance change-point detection methods suffer from p-value bias and inflated Type I error due to data reuse in model selection. Existing post-selection inference (PSI) frameworks are restricted to mean-shift detection and do not extend to variance changes. Method: This paper introduces the first PSI framework for variance change-point detection, proposing two general-purpose constructions for post-selection p-values compatible with diverse algorithms (e.g., piecewise constant modeling) and test forms (e.g., constrained likelihood ratio tests). Leveraging conditional inference, convex optimization, and statistical functional theory, the methods rigorously control Type I error conditional on the selected model path and yield uniformly calibrated p-values. Contribution/Results: We establish theoretical validity of the proposed procedures and demonstrate, via extensive simulations and real-data analyses, their improved statistical power and accurate p-value calibration—overcoming a key limitation of PSI in detecting heteroscedastic structural changes.

Avoiding bias in testing post-selection changepointsExtending post-selection inference to variance changesQuantifying uncertainty in detected variance changepoints

Confidence on the Focal: Conformal Prediction with Selection-Conditional Coverage

Mar 06, 2024
YJ
Ying Jin
🏛️ Harvard University | University of Pennsylvania

Under data-driven selection, conventional prediction intervals fail to guarantee marginal coverage for the selected units—compromising reliability for focal samples. Method: We propose the first finite-sample exact coverage framework for post-selection inference, extending Mondrian conformal prediction to multiple test samples and non-equivariant models while accommodating arbitrary permutation-invariant selection rules. Our approach integrates conditional randomization tests, top-K or optimization-driven selection, conformal p-values, and preliminary screening prediction sets to enable efficient computation. Contribution/Results: Evaluated on drug discovery and health risk prediction tasks, our method substantially improves empirical coverage for focal units, ensuring statistically valid inference in real-world decision-making scenarios. This provides the first provably exact finite-sample coverage guarantee for post-selection prediction intervals under general selection mechanisms.

Addresses selection bias in conformal predictionEnsures valid coverage for selected focal unitsGeneralizes to multiple test units and classifiers

Adaptive A/B Tests and Simultaneous Treatment Parameter Optimization

Oct 13, 2022
YW
Yuhang Wu
🏛️ University of California, Berkeley | Amazon.com Inc

Classical algorithms for strongly convex stochastic optimization achieve fast convergence (O(1/√n)) but suffer from asymptotically non-negligible bias, violating the conditions required for a valid central limit theorem (CLT) and thus impeding asymptotically efficient statistical inference. Method: We propose the first dual-objective algorithm that simultaneously guarantees fast convergence and a provable CLT. Our approach integrates stochastic approximation, asymptotic statistical inference, and adaptive experimental design into a unified framework that ensures asymptotic normality of the estimator. Contribution/Results: We establish theoretical guarantees that the algorithm retains the O(1/√n) convergence rate while satisfying the CLT. Numerical experiments demonstrate substantial improvements over existing methods in estimation accuracy, confidence interval coverage, and identification of optimal treatment parameters. The method provides a new paradigm for continuous, parameterized A/B testing in online platforms—balancing optimization efficiency with statistical reliability.

Addresses non-vanishing bias in stochastic optimization statistical inferenceDevelops algorithm maintaining fast convergence with valid central limit theoremEnables reliable confidence intervals for optimal objective value estimation

Latest Papers

What's happening recently
View more

This work addresses the lack of statistical reliability guarantees in existing hyperparameter selection methods—such as grid search and Bayesian optimization—with respect to critical metrics like risk and safety. Building upon the learn-then-test (LTT) paradigm, the paper introduces a unified statistical framework that formulates hyperparameter selection as a multiple hypothesis testing problem, accommodating user-specified constraints on average risk, quantile risk, or information-theoretic measures. Leveraging tools from statistical inference—including p-values, e-values, and concentration inequalities—the method derives explicit, finite-sample bounds on error probabilities from first principles. This approach enables theoretically grounded validation and selection of hyperparameters, substantially enhancing the reliability and safety of AI systems in real-world deployment scenarios.

hyperparameter selectionmultiple hypothesis testingreliability

This work proposes AutoSI, a novel framework that automates selective inference for any algorithm whose selection event can be expressed as a rational function of the data, eliminating the need for manual derivation by experts. By modeling selection events through rational functions and integrating automatic symbolic computation with exact finite-sample p-value calculation, AutoSI overcomes the limitations of existing methods, which are typically confined to linear or quadratic inequalities. Empirical evaluations across three feature selection tasks—including Lasso tuned via cross-validated R²—demonstrate that AutoSI rigorously controls Type I error while maintaining high statistical power, thereby offering a general, scalable solution for post-selection inference without human intervention.

Feature SelectionRational FunctionsSelection Event

This work addresses Bayesian optimal experimental design under computationally expensive models with limited design evaluations. It proposes an adaptive sequential elimination algorithm that significantly reduces the variance and computational cost of nested Monte Carlo estimators by reusing parameter samples, employing common random numbers, and applying Rao–Blackwellization. A bootstrap-based probabilistic comparison mechanism is integrated to iteratively eliminate inferior designs. The method achieves high reliability while drastically reducing the number of model evaluations, making it well-suited for large-scale engineering applications where computational efficiency and decision accuracy must be carefully balanced.

Bayesian calibrationBayesian optimal experimental designexpensive computational models

This work proposes a Posterior Conformal Selection (PH-CS) framework that overcomes the rigidity of traditional conformal selection methods, which require a pre-specified false discovery rate (FDR) threshold and thus struggle to balance selection size against FDR control. PH-CS eliminates the need for any preset FDR level by constructing a path of candidate selection sets and estimating their data-driven false discovery proportions (FDPs). Leveraging conformal e-values together with the e-BH procedure, the framework enables users to dynamically choose an optimal operating point based on a custom utility function. The method provides reliable average-case FDP estimates under finite samples, extends naturally to general risk control, and demonstrates competitive FDR control performance while accurately estimating FDP and satisfying utility constraints in both synthetic and real-data experiments.

conformal selectione-variablesfalse discovery rate

Hot Scholars

WK

Wonjoong Kim

KAIST
Data MiningLLM AgentQuestion Answering
CP

Chanyoung Park

Associate Professor, KAIST
Artificial intelligenceGraph data miningRecommender systemAI for Science
QH

Qingming Huang

University of the Chinese Academy of Sciences
Multimedia Analysis and RetrievalImage and Video ProcessingPattern RecognitionComputer Vision
JZ

Jiale Zhou

MDH
requirements engineeringsafety critical systemshazard analysisontology
DW

Daniela Witten

Professor of Statistics & Biostatistics, Dorothy Gilford Endowed Chair, University of Washington
statisticsmachine learning