A semi-supervised framework for diverse multiple hypothesis testing scenarios

📅 2024-11-24
📈 Citations: 0
Influential: 0
📄 PDF

career value

169K/year
🤖 AI Summary
In multiple hypothesis testing scenarios where p-values or test statistics are insufficient, conventional methods suffer from low statistical power and lack rigorous finite-sample false discovery rate (FDR) or false discovery proportion (FDP) control. Method: This paper proposes RESET—a novel framework—and its ensemble extension, RESET Ensemble, which for the first time integrates semi-supervised learning with strict finite-sample FDR/FDP control. Leveraging an innovative data-splitting protocol, RESET effectively incorporates side information while preserving theoretical guarantees—without manual hyperparameter tuning or model selection. It is compatible with both p-value–based and competition-based testing paradigms. Contribution/Results: RESET achieves significant power gains over existing methods while maintaining stringent FDR/FDP control under finite samples. It is computationally efficient, theoretically rigorous, and—uniquely among general-purpose multiple testing frameworks—supports user-selectable, provably valid control of either FDR or FDP in finite samples.

Technology Category

Application Category

📝 Abstract
Standard multiple testing procedures are designed to report a list of discoveries, or suspected false null hypotheses, given the hypotheses' p-values or test scores. Recently there has been a growing interest in enhancing such procedures by combining additional information with the primary p-value or score. Specifically, such so-called ``side information'' can be leveraged to improve the separation between true and false nulls along additional ``dimensions'' thereby increasing the overall sensitivity. In line with this idea, we develop RESET (REScoring via Estimating and Training) which uses a unique data-splitting protocol that subsequently allows any semi-supervised learning approach to factor in the available side-information while maintaining finite-sample error rate control. Our practical implementation, RESET Ensemble, selects from an ensemble of classification algorithms so that it is compatible to a range of multiple testing scenarios without the need for the user to select the appropriate one. We apply RESET to both p-value and competition based multiple testing problems and show that RESET is (1) power-wise competitive, (2) fast compared to most tools and (3) is able to uniquely achieve finite sample FDR or FDP control, depending on the user's preference.
Problem

Research questions and friction points this paper is trying to address.

Enhancing multiple testing procedures with side information
Maintaining error rate control in semi-supervised frameworks
Achieving power and speed across diverse testing scenarios
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semi-supervised framework for multiple hypothesis testing
Data-splitting protocol enabling side information integration
Ensemble classification ensuring broad scenario compatibility
🔎 Similar Papers
No similar papers found.