fdr-controlled selection

Design selection procedures that choose a subset of hypotheses, items, or predictions while guaranteeing an upper bound on the expected proportion of false discoveries among the selected set. This involves computing test statistics or calibrated scores (e.g., p-values or conformal scores), choosing data-dependent thresholds and implementing the Benjamini-Hochberg step-up algorithm or equivalent corrections and stability checks to ensure FDR control.

fdr-controlledselection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

False Discovery Rate Adjustments for Average Significance Level Controlling Tests

Sep 27, 2022
TB
Timothy B. Armstrong
🏛️ University of Southern California

Classical false discovery rate (FDR) control methods, such as Benjamini–Hochberg (BH), rely on stringent pointwise control of Type I error (strong control), limiting their applicability under weaker inferential assumptions. This work addresses FDR control when only average-level (i.e., weak) control of significance level is required across tests. Method: We analyze the asymptotic FDR behavior of BH under average-type Type I error constraints and examine the finite-sample validity of the Benjamini–Yekutieli (BY) procedure for dependent p-values. Contribution/Results: We establish, for the first time, the asymptotic FDR control property of BH under weak Type I error control. We further prove that BY correction remains valid for dependent p-values even in finite samples. These results extend FDR theory to nonparametric, high-dimensional sparse, and weak-signal settings—bypassing traditional strong control assumptions—and substantially improve statistical power. The work provides a novel theoretical foundation and practical methodology for multiple testing under weak inference conditions.

Adjusting FDR for tests with average significance level controlEnabling FDR control in nonparametric and high-dimensional settingsExtending BH procedure to weakly dependent p-values asymptotically

More powerful multiple testing under dependence via randomization

May 18, 2023
ZX
Ziyu Xu
🏛️ Carnegie Mellon University

This paper addresses the low statistical power of false discovery rate (FDR) and false coverage rate (FCR) control procedures under arbitrary dependence structures in multiple hypothesis testing. We propose a universal randomization-enhancement strategy based on a single uniform random variable. The method systematically boosts the power of classical procedures—including Benjamini–Yekutieli, e-BH, and Hommel—while preserving exact FDR/FCR control under arbitrary dependence. We provide the first rigorous proof that a single randomization step is never inferior to the original procedure and strictly improves power under any dependence structure; moreover, it unifies and strengthens diverse multiple testing procedures within the e-value framework. Theoretical analysis guarantees strict FDR/FCR control, and extensive simulations confirm substantial power gains. Our core innovation lies in achieving broad-spectrum power enhancement via an extremely simple randomization mechanism, thereby overcoming the long-standing power bottleneck of conventional methods under strong dependence.

Enhances FDR control procedures for dependent p-values and e-valuesImproves power of multiple testing under dependence via randomizationStrengthens Hommel test and post-selection inference for FCR control

E-values, Multiple Testing and Beyond

Dec 05, 2023
GL
Guanxun Li
🏛️ Beijing Normal University at Zhuhai | Texas A&M University

This paper addresses the underutilization of heterogeneity and structural information in multiple hypothesis testing by proposing a general e-value–based framework. Methodologically: (1) it introduces a data-dependent weighting scheme—including a leave-one-out heuristic—for flexible aggregation of e-values across subsets, test statistics, and structure-informed covariates; (2) it unifies and extends the Benjamini–Hochberg (BH) and Benjamini–Yekutieli (BY) procedures to accommodate mixed tests and joint group-level–global false discovery rate (FDR) control; (3) it develops a structure-adaptive e-BH procedure that relaxes the independence and homogeneity assumptions inherent in classical p-value–based methods. Theoretically, it guarantees strict finite-sample FDR control. Numerical experiments demonstrate substantial gains in statistical power over state-of-the-art baselines—particularly under heterogeneous, grouped, or covariate-structured settings.

Controlling false discovery rates in diverse testing scenariosDeveloping e-value aggregation methods for multiple testingIncorporating data-dependent weighting to enhance statistical power

An online generalization of the (e-)Benjamini-Hochberg procedure

Jul 30, 2024
LF
Lasse Fischer
🏛️ University of Bremen | Carnegie Mellon University

Classical Benjamini–Hochberg (BH) and e-BH procedures—designed for offline multiple testing—are not directly applicable to streaming data, and existing online FDR control methods lack rigorous guarantees for stopping-time FDR. Method: We propose an online BH/e-BH framework incorporating a *delayed rejection* mechanism. Contribution/Results: Our method is the first to simultaneously achieve theoretical consistency—provably controlling both fixed-time and stopping-time FDR under arbitrary dependence—and reduction—recovering standard offline BH/e-BH as special cases. We unify and extend the analysis of major online procedures, proving their stopping-time FDR control via sequential decision theory, dependence-agnostic bounding, and stopping-time extension techniques. Crucially, our framework attains the *same* FDR bound as offline BH/e-BH. An open-source implementation enables real-time streaming hypothesis testing.

FDR control for sequential hypothesis testingGeneralizing BH and e-BH methods onlineOnline extension of Benjamini-Hochberg procedure

Anytime-valid FDR control with the stopped e-BH procedure

Feb 12, 2025
HW
Hongjian Wang
🏛️ Carnegie Mellon University

This paper addresses the failure of the stopped e-BH (se-BH) procedure to control the false discovery rate (FDR) under arbitrary stopping times in dependent, data-streaming settings. We identify the root cause as incompatibility between local e-processes and the global filtration, leading to cross-stream information leakage. To resolve this, we introduce verifiable causal conditions—such as absence of unmeasured confounding—and rigorously prove that when e-processes across streams satisfy these conditions, they remain valid e-processes under the global filtration. Consequently, se-BH achieves anytime-valid FDR control for arbitrary stopping times. This is the first online FDR method that simultaneously provides theoretical guarantees and computational feasibility under non-i.i.d., sequentially arriving, and dependent data. The framework enables robust, adaptive multiple testing in real-time applications such as genomics.

Addresses information leakage in adaptively stopped e-processesControls FDR in sequential e-BH procedure with e-processesFormulates causal condition for valid FDR control

Latest Papers

What's happening recently
View more

This study demonstrates that the Benjamini–Hochberg (BH) procedure may fail to control the nominal false discovery rate (FDR) in correlated two-sample Gaussian testing, thereby refuting the long-standing conjecture that BH maintains FDR control under dependence among Gaussian p-values. By constructing a factor model and combining interval arithmetic certification with rigorous mathematical analysis—supplemented by verification using GPT-5.6 Pro—the authors prove that, for sufficiently large numbers of hypotheses and a significance level α = 0.01, the actual FDR strictly exceeds 0.0104. These theoretical findings are corroborated by Monte Carlo simulations, providing the first conclusive evidence that the BH procedure can indeed lose FDR control under specific correlation structures.

Benjamini–Hochberg procedurecorrelated testsfalse discovery rate

This work addresses a limitation of conventional approaches in multiple hypothesis testing within location families, which typically control the false discovery rate (FDR) only under the global null hypothesis. In practice, however, there is often a need to control FDR uniformly over non-significant regions across the entire parameter space. The paper reframes FDR as a function of the location parameter and proposes a natural extension of the Benjamini–Hochberg (BH) procedure that simultaneously controls the entire FDR curve without incurring additional computational cost. Theoretical analysis establishes that the proposed method guarantees the FDR curve remains below a user-specified level uniformly, and numerical experiments corroborate its effectiveness and practical utility.

Benjamini-Hochberg procedurefalse discovery rateFDR control

This study addresses the limited statistical power of existing false discovery rate (FDR) control methods in multiple hypothesis testing by proposing a unified improvement to the Benjamini–Hochberg (BH) procedure based on the e-Closure principle. The proposed method strictly controls the FDR under positive regression dependence on a subset (PRDS) while uniformly dominating the classical BH procedure in power, offering substantial gains when a large proportion of null hypotheses are false. By integrating the e-Closure framework into FDR control for the first time, this work delivers a theoretically rigorous and practically effective enhancement of the BH method, maintaining its robustness while significantly improving detection capability in high-dimensional settings.

Benjamini-Hochberg procedureFalse Discovery Ratemultiple testing

This study investigates the admissibility and complete class problems for false discovery rate (FDR) control procedures within the e-value framework. Drawing on statistical decision theory, it introduces strong and weak dominance relations to establish, for the first time, a theoretical foundation for admissibility in e-value-based multiple testing with FDR control. The main contributions include proving that every step-down procedure is strongly dominated by some weighted average eBH procedure; demonstrating that weighted average eBH procedures without constant terms are admissible at any FDR level; and showing that, under symmetry, this class of procedures forms a complete class, with its members being maximal only when the FDR threshold is sufficiently small—thereby establishing their structural optimality.

admissibilitycomplete classe-values

This work addresses a critical limitation in existing multiple testing procedures, which control only the expected false discovery proportion (FDP) and lack high-probability guarantees for the realized FDP, particularly when data-driven thresholds are employed, thereby compromising statistical validity. The authors propose a distribution-free, finite-sample valid framework that constructs a high-probability simultaneous envelope around the empirical distribution function of conformal p-values under the null hypothesis. This approach yields, for the first time, a uniform high-probability upper bound on the FDP that holds simultaneously over all possible rejection thresholds. The method accommodates arbitrary post-hoc threshold selection and allows users to tailor the envelope’s shape to obtain tighter bounds in regions of interest. Empirical evaluations on both synthetic and real-world data demonstrate that the resulting bounds are not only valid but also substantially less conservative than those from existing methods.

Conformal InferenceDistribution-Free BoundsFalse Discovery Proportion

Hot Scholars

AR

Aaditya Ramdas

Associate Professor (with tenure), Carnegie Mellon University
Machine LearningStatistics
HW

Hongxin Wei

Southern University of Science and Technology (SUSTech)
Reliable Machine LearningUncertainty EstimationStatistics
ED

Edgar Dobriban

Statistics & Computer Science, University of Pennsylvania
StatisticsMachine LearningAI
AS

Armin Schwartzman

Professor, University of California, San Diego
Signal and image analysismanifold-valued datarandom fieldsbrain imaging
SM

Samuel Muller

Executive Dean and Professor, Faculty of Science and Engineering, Macquarie University
StatisticsModel SelectionVariable SelectionRobustness