Asymptotic and compound e-values: multiple testing and empirical Bayes

📅 2024-09-29
📈 Citations: 4
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the theoretical disconnect between false discovery rate (FDR) control and empirical Bayes inference in multiple hypothesis testing. Method: We formally define and systematically develop a rigorous theoretical framework for composite p-values and composite e-values—including asymptotic and fully composite settings—and propose a log-optimal mixture likelihood ratio method for constructing composite e-values. We establish a deep connection between composite e-values and empirical Bayes estimation, prove that any FDR-controlling procedure is equivalent to the e-BH procedure, and derive separable log-optimal composite e-values under point nulls; for heteroscedastic multi-t tests, we construct practical approximate e-values. Contribution: Our work unifies the p-value and e-value literatures, enables e-value derandomization, ensures cross-trial composability, and provides a complete, compositional characterization of FDR control via e-values.

Technology Category

Search and Optimization: Mixed Discrete/Continuous SearchReasoning under Uncertainty: Stochastic OptimizationConstraint Satisfaction and Optimization: Other Foundations of Constraint Satisfaction

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystemsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
📝 Abstract
We explicitly define the notions of (exact, approximate or asymptotic) compound p-values and e-values, which have been implicitly presented and extensively used in the recent multiple testing literature. While it is known that the e-BH procedure with compound e-values controls the FDR, we show the converse: every FDR controlling procedure can be recovered by instantiating the e-BH procedure with certain compound e-values. Since compound e-values are closed under averaging, this allows for combination and derandomization of FDR procedures. We then connect compound e-values to empirical Bayes. In particular, we use the fundamental theorem of compound decision theory to derive the log-optimal simple separable compound e-value for testing a set of point nulls against point alternatives: it is a ratio of mixture likelihoods. We extend universal inference to the compound setting. As one example, we construct approximate compound e-values for multiple t-tests, where the (nuisance) variances may be different across hypotheses. Finally, we provide connections to related notions in the literature stated in terms of p-values.
Problem

Research questions and friction points this paper is trying to address.

Defining compound p-values and e-values for multiple testing
Connecting compound e-values to FDR control procedures
Constructing asymptotic compound e-values for multiple t-tests
Innovation

Methods, ideas, or system contributions that make the work stand out.

Defining compound p-values and e-values
Combining FDR procedures via compound e-values
Constructing asymptotic compound e-values for t-tests
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
University of Chicago | University of Waterloo | Carnegie Mellon University