Score
Designs and implements statistical procedures and software workflows that adjust hypothesis tests for multiplicity, including methods to compute adjusted p-values and to control error rates (e.g., family-wise error rate or false discovery rate) across many comparisons. Builds code and result-table outputs (commonly using R packages) to apply multiple-testing corrections, support paired-comparison tests, and embed pre-specified multiplicity controls into analysis and reporting to prevent inflated false positives.
Existing multiple testing correction methods primarily control Type I error while neglecting Type II error, resulting in suboptimal statistical power—particularly in low-sample-size settings, rare outcomes, or high-dimensional biomarker studies, where detection sensitivity is compromised. To address this, we propose the Beta-Exponent Adjustment (BEA) method, the first framework to explicitly incorporate statistical power into multiple testing correction by jointly optimizing Type I and Type II error rates, thereby balancing false positive control and test sensitivity. Simulation studies under realistic conditions (n = 1000, m = 1000 tests) demonstrate that BEA achieves a sensitivity of 0.8—significantly outperforming classical procedures including Bonferroni, Holm, and Benjamini–Hochberg—while maintaining specificity comparable to these methods. BEA thus establishes a novel paradigm for robust discovery in low-power scenarios.
This study addresses the underappreciated challenge of estimation and communication following multiplicity adjustment within the frequentist framework in complex clinical trials, where multiple endpoints, interim data looks, or group comparisons often introduce estimation bias and complicate interpretation, thereby undermining transparency in benefit–risk assessment. By integrating advanced methodologies such as adaptive designs and graphical approaches to multiple testing, the work illustrates through concrete examples the limitations of current strategies in conveying trial results meaningfully. The research underscores the need to critically reevaluate prevailing practices and foster interdisciplinary dialogue to enhance both the accuracy of effect estimation and the clarity of result communication, ultimately informing future methodological standards and regulatory guidance.
Multiple testing in signal reproducibility detection faces challenges including severe multiple-testing burden and low statistical power in partial conjunction (PC) tests. To address these, we propose a covariate-driven adaptive grouping framework that leverages prior grouping information to enable cross-dataset information sharing, integrates covariate-guided feature screening with hypothesis weighting, and achieves stringent false discovery rate (FDR) control at ≤0.05 in finite samples while substantially improving statistical power. Our method introduces, for the first time, an adaptive filtering mechanism enabling independent weight learning for each hypothesis—overcoming the fundamental power limitations of conventional PC tests. Extensive simulations and analyses of real immune-related gene expression datasets demonstrate that our approach identifies 30–65% more significant reproducible signals than state-of-the-art multilevel testing methods, while rigorously maintaining the target FDR level.
Classical false discovery rate (FDR) control methods, such as Benjamini–Hochberg (BH), rely on stringent pointwise control of Type I error (strong control), limiting their applicability under weaker inferential assumptions. This work addresses FDR control when only average-level (i.e., weak) control of significance level is required across tests. Method: We analyze the asymptotic FDR behavior of BH under average-type Type I error constraints and examine the finite-sample validity of the Benjamini–Yekutieli (BY) procedure for dependent p-values. Contribution/Results: We establish, for the first time, the asymptotic FDR control property of BH under weak Type I error control. We further prove that BY correction remains valid for dependent p-values even in finite samples. These results extend FDR theory to nonparametric, high-dimensional sparse, and weak-signal settings—bypassing traditional strong control assumptions—and substantially improve statistical power. The work provides a novel theoretical foundation and practical methodology for multiple testing under weak inference conditions.
Existing e-BH procedures lack order-invariance over e-processes, causing test conclusions to reverse spuriously upon addition of irrelevant data and failing to control the false discovery rate (FDR) — or even the family-wise error rate (FWER) — under arbitrary dependence. This paper provides the first rigorous proof that e-BH violates FDR control in this setting. Method: We propose a novel, order-invariant multiple testing framework built on e-process upper bounds, featuring a dependence-structure-adaptive calibrator. Contribution/Results: Our method guarantees strict FDR control at level α (i.e., FDR-sup ≤ α) for arbitrary dependence structures among hypotheses. It eliminates temporal instability in rejection sets induced by sequential data arrival, ensuring robustness and reproducibility in dynamic data environments. Theoretical guarantees are established without restrictive assumptions on dependence, and the procedure is computationally tractable.
This study addresses the challenge in multiple hypothesis testing where existing methods struggle with unknown and arbitrary dependence structures among p-values, thereby limiting predictive power analysis and sample size planning. The authors propose the first Bayesian predictive power framework that accommodates arbitrary dependence without requiring independence assumptions, while supporting control of either the family-wise error rate (FWER) or the false discovery rate (FDR). By incorporating prior distributions on effect sizes, a uniform prior on the correlation matrix, and p-value weighting, the approach effectively mitigates p-hacking bias. Inference is carried out via Bayesian simulation using an asymmetric multivariate normal mean-variance mixture distribution with a scale-matrix mixture and a Dirichlet process prior, implemented in the R package bnpMTP. Application to a reanalysis of p-values from a lead exposure study demonstrates more robust power estimation and bias assessment, offering a reliable foundation for future sample size determination.
This study addresses the efficient computation of simultaneous confidence intervals under family-wise error rate (FWER) control in multiple hypothesis testing. By extending, for the first time, the concept of coherence from closed testing procedures to the partitioning principle, the authors establish a formal algorithmic equivalence between these two frameworks, thereby unifying the construction of multiple testing procedures and simultaneous confidence intervals. Leveraging this theoretical connection, they propose a computationally efficient and practically feasible algorithm for constructing simultaneous confidence intervals. The method’s validity and advantages are demonstrated through illustrative examples, highlighting its effectiveness in real-world applications.
In matched case-control studies, conventional statistical analyses of secondary outcomes can yield biased estimates by ignoring the unequal sampling probabilities induced by the matching design. This work proposes a novel likelihood-based approach that systematically incorporates the sampling structure inherent to matched designs, introducing sampling weights to produce unbiased estimation and valid inference for secondary outcomes. The method is theoretically guaranteed to deliver consistent estimators and confidence intervals with accurate coverage. Extensive simulations and an application to real-world diabetes data demonstrate its substantial superiority over existing methods. An R implementation of the proposed approach is publicly available.
This study addresses a critical limitation of traditional multiple testing procedures—such as the Benjamini–Hochberg (BH) method—which control the overall false discovery rate (FDR) but offer no guarantee regarding the reliability of boundary discoveries, i.e., the least significant rejections. The authors propose a novel two-stage adaptive approach: first estimating the number of true null hypotheses using non-significant test statistics, then applying an adjusted threshold within the Support Line (SL) framework to control the error probability of boundary discoveries. This work is the first to integrate adaptivity into boundary FDR control, providing rigorous error guarantees under independence and demonstrating robustness and enhanced power under positive dependence. Theoretical analysis confirms its validity, simulations show substantially improved statistical power over the original SL procedure, and real-world applicability is illustrated through a meta-analysis in psychology.
本文针对高通量数据分析中的多重假设检验问题,提出了一种基于经验累积分布函数的FDP控制方法(eFDP),无需显式建模依赖关系,通过非参数估计和多变量混合模型框架实现更准确的FDP控制。