multiple testing correction

Selecting and applying statistical procedures to combine many hypothesis tests while controlling type I error rates or false discovery (e.g., family-wise error, FDR). Employed to aggregate block- or bin-level evidence, combine repeated tests, and validate power and error control in large-scale experiments.

multipletestingcorrection

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Family-wise error rate control in clinical trials with overlapping populations

Nov 12, 2025
RL
Remi Luschei
🏛️ Competence Center for Clinical Trials Bremen | Institute for Statistics | University of Bremen

In clinical trials with overlapping patient populations, conventional family-wise error rate (FWER) control methods—such as Bonferroni and Holm—that rely on ANOVA-style homogeneity assumptions fail when treatment effects exhibit heterogeneous cancellation across subgroups (e.g., positive in some, negative in others), leading to inflated Type I error. This paper proposes a novel multiple testing correction framework that abandons the homogeneous-effect assumption and instead directly constructs critical values or adjusts α based on the joint distribution of test statistics, explicitly modeling the heterogeneity structure across subgroups. We prove theoretically that the method strictly controls FWER under the generalized null hypothesis. Simulation studies demonstrate substantial improvements in both robustness of error control and statistical power compared to existing approaches. This work establishes a generalizable, heterogeneity-aware paradigm for multiplicity adjustment in precision medicine trials employing overlapping population designs.

Addressing multiplicity issues when testing tailored treatment policies across subgroupsControlling false discoveries in clinical trials with overlapping patient populationsEnsuring family-wise error rate control under heterogeneous null effect scenarios

False Discovery Rate Adjustments for Average Significance Level Controlling Tests

Sep 27, 2022
TB
Timothy B. Armstrong
🏛️ University of Southern California

Classical false discovery rate (FDR) control methods, such as Benjamini–Hochberg (BH), rely on stringent pointwise control of Type I error (strong control), limiting their applicability under weaker inferential assumptions. This work addresses FDR control when only average-level (i.e., weak) control of significance level is required across tests. Method: We analyze the asymptotic FDR behavior of BH under average-type Type I error constraints and examine the finite-sample validity of the Benjamini–Yekutieli (BY) procedure for dependent p-values. Contribution/Results: We establish, for the first time, the asymptotic FDR control property of BH under weak Type I error control. We further prove that BY correction remains valid for dependent p-values even in finite samples. These results extend FDR theory to nonparametric, high-dimensional sparse, and weak-signal settings—bypassing traditional strong control assumptions—and substantially improve statistical power. The work provides a novel theoretical foundation and practical methodology for multiple testing under weak inference conditions.

Adjusting FDR for tests with average significance level controlEnabling FDR control in nonparametric and high-dimensional settingsExtending BH procedure to weakly dependent p-values asymptotically

This study addresses the challenge that conventional variable selection methods struggle to control the false discovery rate (FDR) when predictors are highly correlated and often erroneously discard entire groups of related variables, thereby impairing predictive performance. To overcome this, the authors propose a hierarchical ensemble variable selection framework: variables are first grouped via hierarchical clustering, and each group is then tested for the presence of any non-zero effect, allowing any member to serve as a proxy. This approach uniquely integrates ensemble selection with hierarchical clustering into an FDR-controlling framework—extending beyond prior methods limited to family-wise error rate (FWER) control—and employs a generalized Benjamini–Hochberg/Yekutieli step-up procedure to account for logical dependencies among composite hypotheses. Simulations and empirical analyses demonstrate that the method achieves strict FDR control while substantially improving statistical power, yielding richer and more predictive variable selections.

correlated predictorsfalse discovery ratehierarchical clustering

This study addresses a critical limitation of traditional multiple testing procedures—such as the Benjamini–Hochberg (BH) method—which control the overall false discovery rate (FDR) but offer no guarantee regarding the reliability of boundary discoveries, i.e., the least significant rejections. The authors propose a novel two-stage adaptive approach: first estimating the number of true null hypotheses using non-significant test statistics, then applying an adjusted threshold within the Support Line (SL) framework to control the error probability of boundary discoveries. This work is the first to integrate adaptivity into boundary FDR control, providing rigorous error guarantees under independence and demonstrating robustness and enhanced power under positive dependence. Theoretical analysis confirms its validity, simulations show substantially improved statistical power over the original SL procedure, and real-world applicability is illustrated through a meta-analysis in psychology.

adaptive procedureboundary FDRerror control

Carefree multiple testing with e-processes

Jan 31, 2025
YT
Yury Tavyrikov
🏛️ Vrije Universiteit Amsterdam | Leiden University Medical Centre | University of Twente

Existing e-BH procedures lack order-invariance over e-processes, causing test conclusions to reverse spuriously upon addition of irrelevant data and failing to control the false discovery rate (FDR) — or even the family-wise error rate (FWER) — under arbitrary dependence. This paper provides the first rigorous proof that e-BH violates FDR control in this setting. Method: We propose a novel, order-invariant multiple testing framework built on e-process upper bounds, featuring a dependence-structure-adaptive calibrator. Contribution/Results: Our method guarantees strict FDR control at level α (i.e., FDR-sup ≤ α) for arbitrary dependence structures among hypotheses. It eliminates temporal instability in rejection sets induced by sequential data arrival, ensuring robustness and reproducibility in dynamic data environments. Theoretical guarantees are established without restrictive assumptions on dependence, and the procedure is computationally tractable.

Adjusters needed for FDR control under dependenceCurrent e-value methods fail FDR control with supremaE-processes lack order invariance in multiple testing

Latest Papers

What's happening recently
View more

This study addresses the challenge in multiple hypothesis testing where existing methods struggle with unknown and arbitrary dependence structures among p-values, thereby limiting predictive power analysis and sample size planning. The authors propose the first Bayesian predictive power framework that accommodates arbitrary dependence without requiring independence assumptions, while supporting control of either the family-wise error rate (FWER) or the false discovery rate (FDR). By incorporating prior distributions on effect sizes, a uniform prior on the correlation matrix, and p-value weighting, the approach effectively mitigates p-hacking bias. Inference is carried out via Bayesian simulation using an asymmetric multivariate normal mean-variance mixture distribution with a scale-matrix mixture and a Dirichlet process prior, implemented in the R package bnpMTP. Application to a reanalysis of p-values from a lead exposure study demonstrates more robust power estimation and bias assessment, offering a reliable foundation for future sample size determination.

arbitrary dependencemultiple testing procedurep-value weighting

This work addresses a limitation of conventional approaches in multiple hypothesis testing within location families, which typically control the false discovery rate (FDR) only under the global null hypothesis. In practice, however, there is often a need to control FDR uniformly over non-significant regions across the entire parameter space. The paper reframes FDR as a function of the location parameter and proposes a natural extension of the Benjamini–Hochberg (BH) procedure that simultaneously controls the entire FDR curve without incurring additional computational cost. Theoretical analysis establishes that the proposed method guarantees the FDR curve remains below a user-specified level uniformly, and numerical experiments corroborate its effectiveness and practical utility.

Benjamini-Hochberg procedurefalse discovery rateFDR control

In multiple testing of two-sided Gaussian means, conventional false discovery rate (FDR) procedures—such as the Benjamini–Hochberg (BH) method—fail to control FDR under arbitrary dependence, especially in two-sided settings where dependence structures violate standard assumptions. Method: This paper proposes the first dependency-robust FDR control framework for two-sided tests. It introduces the novel concept of “positive left-tail dependence under the null” (PLTDN), generalizing classical one-sided dependence assumptions to the two-sided case. Based on PLTDN, we construct a family of generalized shift-BH procedures, adaptable to arbitrary covariance structures via p-value adjustment. Contribution/Results: We prove that the proposed method strictly controls FDR under PLTDN. Extensive simulations and analysis of HIV gene expression data demonstrate that, while maintaining FDR ≤ α, it achieves substantially higher statistical power than standard BH—particularly in high-dimensional, strongly correlated settings.

Developing dependence-aware methods for valid false discovery rate controlExtending FDR control to correlated two-sided Gaussian mean testsProposing adjusted p-value procedures that incorporate correlation information

This work aims to enhance the statistical power of adaptive Benjamini–Hochberg (BH) procedures while maintaining control of the false discovery rate (FDR). By unifying existing adaptive FDR methods under a common framework—interpreting them as weighted BH procedures based on composite e-values (ep-BH)—the study reveals their shared structural foundation and demonstrates for the first time that most estimators of the proportion of true null hypotheses inherently correspond to composite e-values. Building on this insight, the authors propose a novel framework that uniformly improves upon nearly all existing methods without requiring additional assumptions, and they develop a new ep-BH procedure with finite-sample FDR guarantees. In canonical settings such as t-tests, the proposed method achieves consistent and robust power gains while rigorously controlling the FDR.

adaptive BH procedurecompound e-valuesfalse discovery rate

Grouped Competition Test with Unified False Discovery Rate Control

Nov 30, 2025
MD
Mingzhou Deng
🏛️ Chinese Academy of Sciences

Existing competitive testing methods struggle to robustly control the false discovery rate (FDR) under strong data heterogeneity or complex dependency structures. To address this, we propose a grouped competitive testing framework that partitions hypotheses according to structural features, designs a calibrated competition mechanism, and achieves rigorous global FDR control via a unified FDR control theorem. Unlike conventional approaches, our method dispenses with the stringent assumption of p-value independence and avoids explicit modeling of dependency structures, thereby balancing flexibility with theoretical rigor. Extensive simulations and real-world mass spectrometry data analyses demonstrate that our method maintains precise FDR control (deviation < 0.5%) under heterogeneous and dependent settings, while significantly outperforming existing competitive tests and BH-type procedures in statistical power—without sacrificing computational efficiency. Our key innovation lies in the first integration of structured hypothesis grouping with calibrated competition, enabling provably valid, low-power-loss FDR control in high-dimensional, heterogeneous data.

Develops a grouped competition test to handle heterogeneous or dependent dataEnsures global False Discovery Rate control with minimal power lossProposes a unified framework for p-value-free multiple hypothesis testing

Hot Scholars

AR

Aaditya Ramdas

Associate Professor (with tenure), Carnegie Mellon University
Machine LearningStatistics
LH

Leonhard Held

Professor of Biostatistics, University of Zurich
StatisticsBiostatisticsEpidemiology
RW

Ruodu Wang

University of Waterloo
StatisticsRisk ManagementActuarial ScienceFinancial Engineering
SP

Samuel Pawel

Epidemiology, Biostatistics and Prevention Institute, University of Zurich
StatisticsMeta-Research