e-value construction

Designing statistical test statistics and e-values with finite-sample validity (including under dependent splits) by adapting concentration inequalities and other techniques to control type I error and produce valid significance measures for hyperparameter selection.

e-valueconstruction

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of statistical reliability guarantees in existing hyperparameter selection methods—such as grid search and Bayesian optimization—with respect to critical metrics like risk and safety. Building upon the learn-then-test (LTT) paradigm, the paper introduces a unified statistical framework that formulates hyperparameter selection as a multiple hypothesis testing problem, accommodating user-specified constraints on average risk, quantile risk, or information-theoretic measures. Leveraging tools from statistical inference—including p-values, e-values, and concentration inequalities—the method derives explicit, finite-sample bounds on error probabilities from first principles. This approach enables theoretically grounded validation and selection of hyperparameters, substantially enhancing the reliability and safety of AI systems in real-world deployment scenarios.

hyperparameter selectionmultiple hypothesis testingreliability

E-values, Multiple Testing and Beyond

Dec 05, 2023
GL
Guanxun Li
🏛️ Beijing Normal University at Zhuhai | Texas A&M University

This paper addresses the underutilization of heterogeneity and structural information in multiple hypothesis testing by proposing a general e-value–based framework. Methodologically: (1) it introduces a data-dependent weighting scheme—including a leave-one-out heuristic—for flexible aggregation of e-values across subsets, test statistics, and structure-informed covariates; (2) it unifies and extends the Benjamini–Hochberg (BH) and Benjamini–Yekutieli (BY) procedures to accommodate mixed tests and joint group-level–global false discovery rate (FDR) control; (3) it develops a structure-adaptive e-BH procedure that relaxes the independence and homogeneity assumptions inherent in classical p-value–based methods. Theoretically, it guarantees strict finite-sample FDR control. Numerical experiments demonstrate substantial gains in statistical power over state-of-the-art baselines—particularly under heterogeneous, grouped, or covariate-structured settings.

Controlling false discovery rates in diverse testing scenariosDeveloping e-value aggregation methods for multiple testingIncorporating data-dependent weighting to enhance statistical power

Revisiting Optimal Allocations for Binary Responses: Insights from Considering Type-I Error Rate Control

Feb 10, 2025
LP
Lukas Pin
🏛️ University of Cambridge | George Mason University

This paper identifies a previously unrecognized inflation of Type I error rate under conventional optimal allocation rules in response-adaptive clinical trials with binary outcomes. To address this, we propose two novel optimal allocation schemes—based on the score test and finite-sample parameter estimation—that jointly optimize statistical power and patient benefit (i.e., minimize expected treatment failures). Our approach avoids Wald-type tests relying on unknown true parameters, thereby enhancing small-sample robustness. Monte Carlo simulations demonstrate that the proposed methods strictly control Type I error while significantly improving patient outcomes—reducing failure rates—in both early-phase and confirmatory trials. The framework naturally extends to multi-arm designs and continuous outcomes, providing both theoretical foundations and practical tools for response-adaptive trial design.

Explores and critiques existing approaches for controlling type-I error ratesInvestigates type-I error rate inflation in optimal response-adaptive designsProposes two new optimal allocation methods for binary outcomes

Carefree multiple testing with e-processes

Jan 31, 2025
YT
Yury Tavyrikov
🏛️ Vrije Universiteit Amsterdam | Leiden University Medical Centre | University of Twente

Existing e-BH procedures lack order-invariance over e-processes, causing test conclusions to reverse spuriously upon addition of irrelevant data and failing to control the false discovery rate (FDR) — or even the family-wise error rate (FWER) — under arbitrary dependence. This paper provides the first rigorous proof that e-BH violates FDR control in this setting. Method: We propose a novel, order-invariant multiple testing framework built on e-process upper bounds, featuring a dependence-structure-adaptive calibrator. Contribution/Results: Our method guarantees strict FDR control at level α (i.e., FDR-sup ≤ α) for arbitrary dependence structures among hypotheses. It eliminates temporal instability in rejection sets induced by sequential data arrival, ensuring robustness and reproducibility in dynamic data environments. Theoretical guarantees are established without restrictive assumptions on dependence, and the procedure is computationally tractable.

Adjusters needed for FDR control under dependenceCurrent e-value methods fail FDR control with supremaE-processes lack order invariance in multiple testing

This study addresses the theoretical gap between classical hypothesis testing with fixed significance levels and Bayesian methods, particularly in light of the Lindley paradox. By leveraging moderate deviation theory, the authors develop a unified Bayesian framework for hypothesis testing. Through Bayesian risk analysis and asymptotic expansions, they show that the optimal test threshold operates on the scale of √(log n / n), naturally yielding Jeffreys’ threshold, the BIC penalty term, and the Chernoff–Stein error exponent. This framework not only resolves the Lindley paradox but also extends Rubin’s (1965) program to modern settings such as high-dimensional sparse inference, goodness-of-fit testing, and model selection. Moreover, it establishes the superiority of Bayesian procedures over classical Neyman–Pearson tests in terms of statistical risk.

Bayes riskBayesian hypothesis testingLindley paradox

Latest Papers

What's happening recently
View more

This study addresses a central challenge in statistical inference for clinical trials: achieving high operational flexibility—such as sample size re-estimation and treatment selection—while rigorously controlling the Type I error rate. The authors systematically integrate confirmatory adaptive designs with e-value-based anytime-valid testing methods, establishing for the first time their formal equivalence through conditional error functions and combination tests, while clarifying their distinct emphases on flexibility. The work constructs a theoretical bridge between these two frameworks: the e-value paradigm enhances optional continuation and loss control, whereas adaptive design principles can refine e-value testing strategies. This synthesis lays both a theoretical foundation and a practical pathway for developing next-generation inferential methods that simultaneously ensure strict error control and substantial procedural adaptability.

anytime-valid testingcombination testsconditional error functions

This work aims to enhance the statistical power of adaptive Benjamini–Hochberg (BH) procedures while maintaining control of the false discovery rate (FDR). By unifying existing adaptive FDR methods under a common framework—interpreting them as weighted BH procedures based on composite e-values (ep-BH)—the study reveals their shared structural foundation and demonstrates for the first time that most estimators of the proportion of true null hypotheses inherently correspond to composite e-values. Building on this insight, the authors propose a novel framework that uniformly improves upon nearly all existing methods without requiring additional assumptions, and they develop a new ep-BH procedure with finite-sample FDR guarantees. In canonical settings such as t-tests, the proposed method achieves consistent and robust power gains while rigorously controlling the FDR.

adaptive BH procedurecompound e-valuesfalse discovery rate

This work proposes an e-value–based adaptive design for small-sample, single-arm clinical trials with binary outcomes, enabling flexible interim analyses and automatic futility stopping while rigorously controlling Type I error. It is the first to apply a finite-horizon optimal e-value framework to this setting, leveraging a betting interpretation combined with constrained dynamic programming to optimize either finite-sample statistical power or expected sample size. Compared to conventional approaches and asymptotically optimal e-value designs, the proposed method demonstrates superior performance in small samples, offering both robustness and efficiency. Moreover, it facilitates real-time futility detection through extremely small e-values, allowing early termination when treatment efficacy is implausible.

adaptive clinical trialsbinary datae-values

This study addresses the challenge of constructing valid confidence intervals in two-stage adaptive enrichment clinical trials, where patient subgroups are selected based on interim data, thereby compromising the nominal coverage of conventional intervals. The authors propose a novel method that constructs confidence intervals conditional on the interim selection decision, leveraging conditional inference and inversion of uniformly most accurate unbiased (UMAU) tests to guarantee exact coverage within the selected subgroup. The approach is broadly applicable to various adaptive enrichment designs and is implemented via an efficient numerical algorithm. Extensive simulation studies demonstrate that the proposed intervals consistently achieve the desired coverage probability across diverse design configurations, substantially outperforming existing methods in both validity and precision.

adaptive enrichment designsconfidence intervalscoverage probability

This study addresses the challenge of effectively combining e-values under data-dependent tuning while preserving statistical validity. The authors develop a class of optimized e-process–based combination methods tailored for both independent settings and a newly introduced notion of “simultaneous e-variables.” They establish, for the first time, that combined tests retain validity even when tuning parameters are selected in a data-driven manner. By incorporating elementary symmetric polynomials, the proposed approach enhances statistical power and establishes a flexible intermediate framework that bridges independent and sequential validity. Through rigorous theoretical analysis grounded in dependence structure modeling, the method achieves substantially improved detection capability while maintaining strict Type-I error control.

combination testsdata-dependent tuninge-values

Hot Scholars

AR

Aaditya Ramdas

Associate Professor (with tenure), Carnegie Mellon University
Machine LearningStatistics
YF

Yebo Feng

Nanyang Technological University
Computer SecurityNetwork SecurityBlockchain SecurityNetwork Traffic Analysis
MS

Michael S. Bernstein

Professor of Computer Science, Stanford University
Human-computer interactionsocial computinghuman-centered AI
ZE

Ziv Epstein

MIT
computational social sciencesocial mediaartificial intelligence
FJ

Farnaz Jahanbakhsh

University of Michigan
Human-Computer InteractionSocial Computing