🤖 AI Summary
Existing methods for assessing cross-study replicability of high-dimensional data suffer from strong parametric assumptions, poor scalability to multi-study settings, and insufficient statistical power. To address these limitations, this paper proposes the first empirical Bayes framework that simultaneously achieves high statistical power and strong robustness. Our method adaptively integrates information across multiple features and studies, explicitly models effect heterogeneity, and guarantees strict false discovery rate (FDR) control via theoretical analysis. Applied to real genome-wide association study (GWAS) data, it substantially improves detection of replicable signals and uncovers several biologically meaningful findings missed by mainstream approaches. Both theoretical analysis and empirical evaluation demonstrate that our method consistently outperforms state-of-the-art techniques in statistical power, robustness to model misspecification, and interpretability.
📝 Abstract
Identifying replicable signals across different studies provides stronger scientific evidence and more powerful inference. Existing literature on high dimensional applicability analysis either imposes strong modeling assumptions or has low power. We develop a powerful and robust empirical Bayes approach for high dimensional replicability analysis. Our method effectively borrows information from different features and studies while accounting for heterogeneity. We show that the proposed method has better power than competing methods while controlling the false discovery rate, both empirically and theoretically. Analyzing datasets from the genome-wide association studies reveals new biological insights that otherwise cannot be obtained by using existing methods.