family-wise error control

Design, implement, and analyze statistical procedures that control the family-wise error rate (FWER) — the probability of making one or more false rejections across a family of hypothesis tests. This includes adjusting significance thresholds and applying or evaluating corrections such as Bonferroni and Holm, and assessing their error control, conservativeness, and power trade-offs.

family-wiseerrorcontrol

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Family-wise error rate control in clinical trials with overlapping populations

Nov 12, 2025
RL
Remi Luschei
🏛️ Competence Center for Clinical Trials Bremen | Institute for Statistics | University of Bremen

In clinical trials with overlapping patient populations, conventional family-wise error rate (FWER) control methods—such as Bonferroni and Holm—that rely on ANOVA-style homogeneity assumptions fail when treatment effects exhibit heterogeneous cancellation across subgroups (e.g., positive in some, negative in others), leading to inflated Type I error. This paper proposes a novel multiple testing correction framework that abandons the homogeneous-effect assumption and instead directly constructs critical values or adjusts α based on the joint distribution of test statistics, explicitly modeling the heterogeneity structure across subgroups. We prove theoretically that the method strictly controls FWER under the generalized null hypothesis. Simulation studies demonstrate substantial improvements in both robustness of error control and statistical power compared to existing approaches. This work establishes a generalizable, heterogeneity-aware paradigm for multiplicity adjustment in precision medicine trials employing overlapping population designs.

Addressing multiplicity issues when testing tailored treatment policies across subgroupsControlling false discoveries in clinical trials with overlapping patient populationsEnsuring family-wise error rate control under heterogeneous null effect scenarios

This study addresses the vulnerability of traditional population-wise error rate (PWER) methods in overlapping subgroup clinical trials for personalized medicine, where Type I error control for a specific subgroup can be adversely affected by other subgroups. To overcome this limitation, the authors propose two novel multiple error rate control procedures—PWER-P and PWER-U—that enforce individualized error rate control over all target subgroups and any arbitrary union of them, respectively. Built upon probability theory and multiple hypothesis testing principles, the proposed framework rigorously controls the Type I error while substantially enhancing statistical power. Compared with conventional PWER and family-wise error rate (FWER) approaches, these new methods demonstrate superior stability, controllability, and practical utility in settings involving overlapping subpopulations.

clinical trialserror rate controlmultiple type I error

Family-wise Error Rate Control with E-values

Jan 15, 2025
WH
Will Hartog
🏛️ Stanford University

This paper addresses the challenge of controlling the family-wise error rate (FWER) in multiple hypothesis testing under the e-value framework. We propose the first graph-based closed testing procedure grounded in e-values. By extending graphical methods to the e-value domain—integrating directed acyclic graph (DAG) modeling with dynamic programming—we design a polynomial-time closed testing algorithm. Our approach achieves near-linear time complexity for Holm’s procedure and general fallback graphs: worst-case O(m²), and O(m log m) for specific DAGs. Crucially, it operates directly on e-values without conversion to p-values, thereby substantially improving statistical power. The method accommodates arbitrary DAG structures, overcoming both the computational bottlenecks and power limitations inherent in existing e-value-based closed testing procedures.

Develops efficient algorithms for e-value-based approaches using dynamic programmingEnhances computational efficiency for special graphs like Holm's procedureExtends graphical FWER control to e-values for more powerful testing

This study addresses the challenge of multiple testing in online change-point detection, where repeated hypothesis tests render traditional family-wise error rate (FWER) control inadequate. The work introduces the first formal definition of sequential FWER (sFWER) tailored to real-time data streams and proposes a simulation-based dynamic threshold calibration mechanism that effectively controls the probability of at least one false alarm within a sliding monitoring window. By integrating sequential inference with a moving window framework, the method operates without requiring a pre-specified number of tests and is well-suited for highly dependent data. Empirical evaluations demonstrate that the proposed approach accurately controls sFWER and substantially outperforms existing methods, with successful application to passive smartphone sensing data from adolescents exhibiting emotional instability.

false positivesfamily-wise error ratemultiple testing

False Discovery Rate Adjustments for Average Significance Level Controlling Tests

Sep 27, 2022
TB
Timothy B. Armstrong
🏛️ University of Southern California

Classical false discovery rate (FDR) control methods, such as Benjamini–Hochberg (BH), rely on stringent pointwise control of Type I error (strong control), limiting their applicability under weaker inferential assumptions. This work addresses FDR control when only average-level (i.e., weak) control of significance level is required across tests. Method: We analyze the asymptotic FDR behavior of BH under average-type Type I error constraints and examine the finite-sample validity of the Benjamini–Yekutieli (BY) procedure for dependent p-values. Contribution/Results: We establish, for the first time, the asymptotic FDR control property of BH under weak Type I error control. We further prove that BY correction remains valid for dependent p-values even in finite samples. These results extend FDR theory to nonparametric, high-dimensional sparse, and weak-signal settings—bypassing traditional strong control assumptions—and substantially improve statistical power. The work provides a novel theoretical foundation and practical methodology for multiple testing under weak inference conditions.

Adjusting FDR for tests with average significance level controlEnabling FDR control in nonparametric and high-dimensional settingsExtending BH procedure to weakly dependent p-values asymptotically

Latest Papers

What's happening recently
View more

This study addresses the challenge of enhancing statistical power while controlling the family-wise error rate (FWER) in multiple hypothesis testing when hypotheses carry unequal importance. The authors systematically compare two weighted Holm procedures: the Weighted Holm based on weighted p-value ordering (WHP) and the Weighted Holm based on unweighted p-value ordering (WAP). Through closure principles, graphical representations, and adjusted p-value derivations, they theoretically demonstrate that WHP uniformly dominates WAP—WHP satisfies both monotonicity and admissibility, whereas WAP, although consonant, lacks monotonicity and is optimal only under restrictive conditions. Both theoretical analysis and simulations confirm that WHP achieves superior FWER control and higher average power, offering a clear recommendation for practical implementation.

familywise error rateFWER controlhypothesis weighting

This study addresses the loss of statistical power in multiple hypothesis testing due to reliance solely on p-value ordering. Focusing on K exchangeable hypotheses, the authors propose an optimal family-wise error rate (FWER)–controlling procedure based on elementary symmetric polynomials of likelihood ratios. The method maximizes statistical power while strictly controlling the FWER. Key contributions include the first computable closed-form solution for the dual vector, a proof of global monotonicity of the objective function ensuring a unique coordinate root, and an efficient, scalable algorithm combining coordinate descent with bisection search. Empirical results demonstrate substantial power gains over Hommel’s method—15% at K=3 and 84% at K=12—with superior performance validated in replication studies and clinical trials.

exchangeable hypothesesfamily-wise error ratemultiple testing

This study addresses the challenge that conventional variable selection methods struggle to control the false discovery rate (FDR) when predictors are highly correlated and often erroneously discard entire groups of related variables, thereby impairing predictive performance. To overcome this, the authors propose a hierarchical ensemble variable selection framework: variables are first grouped via hierarchical clustering, and each group is then tested for the presence of any non-zero effect, allowing any member to serve as a proxy. This approach uniquely integrates ensemble selection with hierarchical clustering into an FDR-controlling framework—extending beyond prior methods limited to family-wise error rate (FWER) control—and employs a generalized Benjamini–Hochberg/Yekutieli step-up procedure to account for logical dependencies among composite hypotheses. Simulations and empirical analyses demonstrate that the method achieves strict FDR control while substantially improving statistical power, yielding richer and more predictive variable selections.

correlated predictorsfalse discovery ratehierarchical clustering

This study addresses the efficient computation of simultaneous confidence intervals under family-wise error rate (FWER) control in multiple hypothesis testing. By extending, for the first time, the concept of coherence from closed testing procedures to the partitioning principle, the authors establish a formal algorithmic equivalence between these two frameworks, thereby unifying the construction of multiple testing procedures and simultaneous confidence intervals. Leveraging this theoretical connection, they propose a computationally efficient and practically feasible algorithm for constructing simultaneous confidence intervals. The method’s validity and advantages are demonstrated through illustrative examples, highlighting its effectiveness in real-world applications.

closed testingcomputational efficiencyfamily-wise error rate

This study addresses the challenge of accurately identifying subregions with genuine treatment effects in multi-site randomized policy experiments, where conventional multiple testing procedures often lack power. The authors propose a top-down, tree-structured sequential testing procedure that begins by evaluating the overall effect and then recursively tests groups of sites and individual sites, halting further testing within any branch once non-significance is encountered. This approach innovatively integrates a hierarchical tree framework, Hommel-type correction, and adaptive α allocation to enhance detection power for heterogeneous effects while rigorously controlling the weak familywise error rate (FWER). In simulations based on a 44-site education experiment, the method achieved a 44% detection rate—approximately four times higher than the standard Hommel procedure (11%)—and was successfully applied across 25 MDRC education trials.

block-level effectsfamily-wise error rateheterogeneous treatment effects

Hot Scholars

AV

Anna Vesely

Dep. of Statistical Sciences, University of Bologna
multiple hypothesis testingselective inferencepermutation testinghigh-dimensional data
AA

Angela Andreella

Ca' Foscari University of Venice
Multivariate analysisSocial StatisticsPsychometricsHigh-dimensional data
EX

Ethan X. Fang

Associate Professor at Duke University
StatisticsBiostatisticsOptimization
HG

Hongfu Gao

National University of Singapore
Reliable Machine LearningNatural Language Processing
BJ

Bingyi Jing

Chair Professor, Southern University of Science & Technology
StatisticsData ScienceAI