Testing for Replicated Signals Across Multiple Studies with Side Information

📅 2025-05-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Multiple testing in signal reproducibility detection faces challenges including severe multiple-testing burden and low statistical power in partial conjunction (PC) tests. To address these, we propose a covariate-driven adaptive grouping framework that leverages prior grouping information to enable cross-dataset information sharing, integrates covariate-guided feature screening with hypothesis weighting, and achieves stringent false discovery rate (FDR) control at ≤0.05 in finite samples while substantially improving statistical power. Our method introduces, for the first time, an adaptive filtering mechanism enabling independent weight learning for each hypothesis—overcoming the fundamental power limitations of conventional PC tests. Extensive simulations and analyses of real immune-related gene expression datasets demonstrate that our approach identifies 30–65% more significant reproducible signals than state-of-the-art multilevel testing methods, while rigorously maintaining the target FDR level.

Technology Category

Machine Learning: Multi-instance/Multi-view LearningSearch and Optimization: Mixed Discrete/Continuous SearchKnowledge Representation and Reasoning: Diagnosis and Abductive Reasoning

Application Category

Graph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
Partial conjunction (PC) $p$-values and side information provided by covariates can be used to detect signals that replicate across multiple studies investigating the same set of features, all while controlling the false discovery rate (FDR). However, when many features are present, the extent of multiplicity correction required for FDR control, along with the inherently limited power of PC $p$-values$unicode{x2013}$especially when replication across all studies is demanded$unicode{x2013}$often inhibits the number of discoveries made. To address this problem, we develop a $p$-value-based covariate-adaptive methodology that revolves around partitioning studies into smaller groups and borrowing information between them to filter out unpromising features. This filtering strategy: 1) reduces the multiplicity correction required for FDR control, and 2) allows independent hypothesis weights to be trained on data from filtered-out features to enhance the power of the PC $p$-values in the rejection rule. Our methodology has finite-sample FDR control under minimal distributional assumptions, and we demonstrate its competitive performance through simulation studies and a real-world case study on gene expression and the immune system.
Problem

Research questions and friction points this paper is trying to address.

Detect replicated signals across studies using covariates and PC p-values
Address low discovery power due to stringent multiplicity correction
Enhance signal detection by partitioning studies and borrowing information
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses partial conjunction p-values with covariates
Partitions studies into smaller adaptive groups
Trains hypothesis weights on filtered features
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Ninh Tran
D
Dennis Leung