multiple comparisons correction

Designs and implements statistical procedures and software workflows that adjust hypothesis tests for multiplicity, including methods to compute adjusted p-values and to control error rates (e.g., family-wise error rate or false discovery rate) across many comparisons. Builds code and result-table outputs (commonly using R packages) to apply multiple-testing corrections, support paired-comparison tests, and embed pre-specified multiplicity controls into analysis and reporting to prevent inflated false positives.

multiplecomparisonscorrection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

An alternative method of adjusting for multiple comparison in medical research

Jul 30, 2025
JL
Jiale Li
🏛️ Southern Medical University | Tianjin Nankai High School

Existing multiple testing correction methods primarily control Type I error while neglecting Type II error, resulting in suboptimal statistical power—particularly in low-sample-size settings, rare outcomes, or high-dimensional biomarker studies, where detection sensitivity is compromised. To address this, we propose the Beta-Exponent Adjustment (BEA) method, the first framework to explicitly incorporate statistical power into multiple testing correction by jointly optimizing Type I and Type II error rates, thereby balancing false positive control and test sensitivity. Simulation studies under realistic conditions (n = 1000, m = 1000 tests) demonstrate that BEA achieves a sensitivity of 0.8—significantly outperforming classical procedures including Bonferroni, Holm, and Benjamini–Hochberg—while maintaining specificity comparable to these methods. BEA thus establishes a novel paradigm for robust discovery in low-power scenarios.

Compares BEA with Bonferroni, Holm, and BH methodsEnhances sensitivity and power in multiple testing correctionsProposes BEA method to control Type I and II errors

This study addresses the underappreciated challenge of estimation and communication following multiplicity adjustment within the frequentist framework in complex clinical trials, where multiple endpoints, interim data looks, or group comparisons often introduce estimation bias and complicate interpretation, thereby undermining transparency in benefit–risk assessment. By integrating advanced methodologies such as adaptive designs and graphical approaches to multiple testing, the work illustrates through concrete examples the limitations of current strategies in conveying trial results meaningfully. The research underscores the need to critically reevaluate prevailing practices and foster interdisciplinary dialogue to enhance both the accuracy of effect estimation and the clarity of result communication, ultimately informing future methodological standards and regulatory guidance.

adaptive designestimationmultiple hypotheses

Multiple testing in signal reproducibility detection faces challenges including severe multiple-testing burden and low statistical power in partial conjunction (PC) tests. To address these, we propose a covariate-driven adaptive grouping framework that leverages prior grouping information to enable cross-dataset information sharing, integrates covariate-guided feature screening with hypothesis weighting, and achieves stringent false discovery rate (FDR) control at ≤0.05 in finite samples while substantially improving statistical power. Our method introduces, for the first time, an adaptive filtering mechanism enabling independent weight learning for each hypothesis—overcoming the fundamental power limitations of conventional PC tests. Extensive simulations and analyses of real immune-related gene expression datasets demonstrate that our approach identifies 30–65% more significant reproducible signals than state-of-the-art multilevel testing methods, while rigorously maintaining the target FDR level.

Address low discovery power due to stringent multiplicity correctionDetect replicated signals across studies using covariates and PC p-valuesEnhance signal detection by partitioning studies and borrowing information

False Discovery Rate Adjustments for Average Significance Level Controlling Tests

Sep 27, 2022
TB
Timothy B. Armstrong
🏛️ University of Southern California

Classical false discovery rate (FDR) control methods, such as Benjamini–Hochberg (BH), rely on stringent pointwise control of Type I error (strong control), limiting their applicability under weaker inferential assumptions. This work addresses FDR control when only average-level (i.e., weak) control of significance level is required across tests. Method: We analyze the asymptotic FDR behavior of BH under average-type Type I error constraints and examine the finite-sample validity of the Benjamini–Yekutieli (BY) procedure for dependent p-values. Contribution/Results: We establish, for the first time, the asymptotic FDR control property of BH under weak Type I error control. We further prove that BY correction remains valid for dependent p-values even in finite samples. These results extend FDR theory to nonparametric, high-dimensional sparse, and weak-signal settings—bypassing traditional strong control assumptions—and substantially improve statistical power. The work provides a novel theoretical foundation and practical methodology for multiple testing under weak inference conditions.

Adjusting FDR for tests with average significance level controlEnabling FDR control in nonparametric and high-dimensional settingsExtending BH procedure to weakly dependent p-values asymptotically

Carefree multiple testing with e-processes

Jan 31, 2025
YT
Yury Tavyrikov
🏛️ Vrije Universiteit Amsterdam | Leiden University Medical Centre | University of Twente

Existing e-BH procedures lack order-invariance over e-processes, causing test conclusions to reverse spuriously upon addition of irrelevant data and failing to control the false discovery rate (FDR) — or even the family-wise error rate (FWER) — under arbitrary dependence. This paper provides the first rigorous proof that e-BH violates FDR control in this setting. Method: We propose a novel, order-invariant multiple testing framework built on e-process upper bounds, featuring a dependence-structure-adaptive calibrator. Contribution/Results: Our method guarantees strict FDR control at level α (i.e., FDR-sup ≤ α) for arbitrary dependence structures among hypotheses. It eliminates temporal instability in rejection sets induced by sequential data arrival, ensuring robustness and reproducibility in dynamic data environments. Theoretical guarantees are established without restrictive assumptions on dependence, and the procedure is computationally tractable.

Adjusters needed for FDR control under dependenceCurrent e-value methods fail FDR control with supremaE-processes lack order invariance in multiple testing

Latest Papers

What's happening recently
View more

This study addresses the challenge in multiple hypothesis testing where existing methods struggle with unknown and arbitrary dependence structures among p-values, thereby limiting predictive power analysis and sample size planning. The authors propose the first Bayesian predictive power framework that accommodates arbitrary dependence without requiring independence assumptions, while supporting control of either the family-wise error rate (FWER) or the false discovery rate (FDR). By incorporating prior distributions on effect sizes, a uniform prior on the correlation matrix, and p-value weighting, the approach effectively mitigates p-hacking bias. Inference is carried out via Bayesian simulation using an asymmetric multivariate normal mean-variance mixture distribution with a scale-matrix mixture and a Dirichlet process prior, implemented in the R package bnpMTP. Application to a reanalysis of p-values from a lead exposure study demonstrates more robust power estimation and bias assessment, offering a reliable foundation for future sample size determination.

arbitrary dependencemultiple testing procedurep-value weighting

This study addresses the efficient computation of simultaneous confidence intervals under family-wise error rate (FWER) control in multiple hypothesis testing. By extending, for the first time, the concept of coherence from closed testing procedures to the partitioning principle, the authors establish a formal algorithmic equivalence between these two frameworks, thereby unifying the construction of multiple testing procedures and simultaneous confidence intervals. Leveraging this theoretical connection, they propose a computationally efficient and practically feasible algorithm for constructing simultaneous confidence intervals. The method’s validity and advantages are demonstrated through illustrative examples, highlighting its effectiveness in real-world applications.

closed testingcomputational efficiencyfamily-wise error rate

In matched case-control studies, conventional statistical analyses of secondary outcomes can yield biased estimates by ignoring the unequal sampling probabilities induced by the matching design. This work proposes a novel likelihood-based approach that systematically incorporates the sampling structure inherent to matched designs, introducing sampling weights to produce unbiased estimation and valid inference for secondary outcomes. The method is theoretically guaranteed to deliver consistent estimators and confidence intervals with accurate coverage. Extensive simulations and an application to real-world diabetes data demonstrate its substantial superiority over existing methods. An R implementation of the proposed approach is publicly available.

epidemiological researchmatched case-control studiessecondary outcomes

This study addresses a critical limitation of traditional multiple testing procedures—such as the Benjamini–Hochberg (BH) method—which control the overall false discovery rate (FDR) but offer no guarantee regarding the reliability of boundary discoveries, i.e., the least significant rejections. The authors propose a novel two-stage adaptive approach: first estimating the number of true null hypotheses using non-significant test statistics, then applying an adjusted threshold within the Support Line (SL) framework to control the error probability of boundary discoveries. This work is the first to integrate adaptivity into boundary FDR control, providing rigorous error guarantees under independence and demonstrating robustness and enhanced power under positive dependence. Theoretical analysis confirms its validity, simulations show substantially improved statistical power over the original SL procedure, and real-world applicability is illustrated through a meta-analysis in psychology.

adaptive procedureboundary FDRerror control

Hot Scholars

HW

Hongxin Wei

Southern University of Science and Technology (SUSTech)
Reliable Machine LearningUncertainty EstimationStatistics
OS

Osvaldo Simeone

King's College London
Information theorymachine learningquantum information processingwireless systems
WS

Wenguang Sun

Professor of Data Sciences and Operations, University of Southern California
Large-scale Multiple TestingDecision TheoryHigh Dimensional Statistical Inference
SM

Shima Mohammadi

PhD student, Instituto Superior Técnico
Image CompressionVisual Quality AssessmentMultimedia Signal Processing