Score
Designs and analyzes estimation and uncertainty-quantification procedures that remain valid after data-driven selection of models, features, or hyperparameters, producing selection-adjusted confidence intervals, p-values, and tuning rules that control selection-induced bias and error probabilities. This includes building adaptive hyperparameter-selection algorithms and calibration tests (e.g., learn-then-test, multiple-testing corrections, aggregation-of-candidate-tests, jackknife or other resampling calibrations) that guarantee conservative or finite-sample coverage and meet user-specified reliability constraints.
Traditional statistical inference often fails when models or parameters are selected in a data-driven manner. This work systematically investigates selective inference, focusing on a conditional inference framework that conditions on the selection event to restore the nominal coverage of confidence intervals. We clarify the scientific interpretation of this framework, unify several existing approaches under its umbrella, and demonstrate its application to canonical settings such as inference for the “winner,” region-specific means in regression trees, and differences between clusters. Through simulations and analyses of single-cell RNA sequencing data, we show that the proposed methodology yields valid and reliable statistical inference in practical scenarios.
In high-dimensional personalized treatment strategy estimation, standard post-variable-selection statistical inference fails due to selection-induced bias. Method: This paper introduces Universal Post-Selection Inference (UPoSI) into the robust Q-learning framework, proposing a selection-mechanism-agnostic universal post-selection inference method. The approach uniformly improves confidence interval construction and is theoretically shown to be asymptotically valid in multi-stage decision settings, guaranteeing nominal Type-I error control for hypothesis testing and exact coverage probability for confidence intervals. Results: Monte Carlo simulations demonstrate that the proposed method substantially improves coverage accuracy and statistical power compared to selective inference, while remaining compatible with diverse data-driven variable selection procedures. It thus provides a generalizable and verifiable foundation for statistical inference in robust Q-learning.
In high-cost or high-risk settings where hyperparameter selection is constrained by limited test budgets and demanding statistical reliability requirements, this paper proposes the adaptive Learn-then-Test (aLTT) framework. Methodologically, aLTT introduces e-process theory—novelly applied to sequential, data-dependent multiple hypothesis testing—to enable rigorous false discovery rate (FDR) control with provably valid early stopping. Evaluated on offline reinforcement learning policy selection and large language model prompt optimization, aLTT achieves comparable statistical guarantees and final performance to classical LTT, while reducing the number of required test rounds by an order of magnitude. By unifying finite-sample statistical rigor with practical engineering efficiency, aLTT establishes a provably reliable, sample-efficient paradigm for AI model evaluation under resource constraints.
Classical variance change-point detection methods suffer from p-value bias and inflated Type I error due to data reuse in model selection. Existing post-selection inference (PSI) frameworks are restricted to mean-shift detection and do not extend to variance changes. Method: This paper introduces the first PSI framework for variance change-point detection, proposing two general-purpose constructions for post-selection p-values compatible with diverse algorithms (e.g., piecewise constant modeling) and test forms (e.g., constrained likelihood ratio tests). Leveraging conditional inference, convex optimization, and statistical functional theory, the methods rigorously control Type I error conditional on the selected model path and yield uniformly calibrated p-values. Contribution/Results: We establish theoretical validity of the proposed procedures and demonstrate, via extensive simulations and real-data analyses, their improved statistical power and accurate p-value calibration—overcoming a key limitation of PSI in detecting heteroscedastic structural changes.
Under data-driven selection, conventional prediction intervals fail to guarantee marginal coverage for the selected units—compromising reliability for focal samples. Method: We propose the first finite-sample exact coverage framework for post-selection inference, extending Mondrian conformal prediction to multiple test samples and non-equivariant models while accommodating arbitrary permutation-invariant selection rules. Our approach integrates conditional randomization tests, top-K or optimization-driven selection, conformal p-values, and preliminary screening prediction sets to enable efficient computation. Contribution/Results: Evaluated on drug discovery and health risk prediction tasks, our method substantially improves empirical coverage for focal units, ensuring statistically valid inference in real-world decision-making scenarios. This provides the first provably exact finite-sample coverage guarantee for post-selection prediction intervals under general selection mechanisms.
Classical algorithms for strongly convex stochastic optimization achieve fast convergence (O(1/√n)) but suffer from asymptotically non-negligible bias, violating the conditions required for a valid central limit theorem (CLT) and thus impeding asymptotically efficient statistical inference. Method: We propose the first dual-objective algorithm that simultaneously guarantees fast convergence and a provable CLT. Our approach integrates stochastic approximation, asymptotic statistical inference, and adaptive experimental design into a unified framework that ensures asymptotic normality of the estimator. Contribution/Results: We establish theoretical guarantees that the algorithm retains the O(1/√n) convergence rate while satisfying the CLT. Numerical experiments demonstrate substantial improvements over existing methods in estimation accuracy, confidence interval coverage, and identification of optimal treatment parameters. The method provides a new paradigm for continuous, parameterized A/B testing in online platforms—balancing optimization efficiency with statistical reliability.
This work addresses the lack of statistical reliability guarantees in existing hyperparameter selection methods—such as grid search and Bayesian optimization—with respect to critical metrics like risk and safety. Building upon the learn-then-test (LTT) paradigm, the paper introduces a unified statistical framework that formulates hyperparameter selection as a multiple hypothesis testing problem, accommodating user-specified constraints on average risk, quantile risk, or information-theoretic measures. Leveraging tools from statistical inference—including p-values, e-values, and concentration inequalities—the method derives explicit, finite-sample bounds on error probabilities from first principles. This approach enables theoretically grounded validation and selection of hyperparameters, substantially enhancing the reliability and safety of AI systems in real-world deployment scenarios.
This work proposes AutoSI, a novel framework that automates selective inference for any algorithm whose selection event can be expressed as a rational function of the data, eliminating the need for manual derivation by experts. By modeling selection events through rational functions and integrating automatic symbolic computation with exact finite-sample p-value calculation, AutoSI overcomes the limitations of existing methods, which are typically confined to linear or quadratic inequalities. Empirical evaluations across three feature selection tasks—including Lasso tuned via cross-validated R²—demonstrate that AutoSI rigorously controls Type I error while maintaining high statistical power, thereby offering a general, scalable solution for post-selection inference without human intervention.
This work addresses Bayesian optimal experimental design under computationally expensive models with limited design evaluations. It proposes an adaptive sequential elimination algorithm that significantly reduces the variance and computational cost of nested Monte Carlo estimators by reusing parameter samples, employing common random numbers, and applying Rao–Blackwellization. A bootstrap-based probabilistic comparison mechanism is integrated to iteratively eliminate inferior designs. The method achieves high reliability while drastically reducing the number of model evaluations, making it well-suited for large-scale engineering applications where computational efficiency and decision accuracy must be carefully balanced.
This work proposes a Posterior Conformal Selection (PH-CS) framework that overcomes the rigidity of traditional conformal selection methods, which require a pre-specified false discovery rate (FDR) threshold and thus struggle to balance selection size against FDR control. PH-CS eliminates the need for any preset FDR level by constructing a path of candidate selection sets and estimating their data-driven false discovery proportions (FDPs). Leveraging conformal e-values together with the e-BH procedure, the framework enables users to dynamically choose an optimal operating point based on a custom utility function. The method provides reliable average-case FDP estimates under finite samples, extends naturally to general risk control, and demonstrates competitive FDR control performance while accurately estimating FDP and satisfying utility constraints in both synthetic and real-data experiments.