Active Hypothesis Testing under Computational Budgets with Applications to GWAS and LLM

📅 2025-12-01
📈 Citations: 0
Influential: 0
📄 PDF

career value

203K/year
🤖 AI Summary
To address the computational budget limitations in large-scale hypothesis testing—where exact $p$- or $e$-value computation is often infeasible—this paper proposes a budget-aware active testing framework. The method leverages auxiliary statistics to adaptively decide, via a probabilistic mechanism, whether to compute exact or efficient surrogate test statistics, ensuring strict budget adherence in expectation. We establish theoretical guarantees of statistical optimality and validity under both independence and dependence assumptions. The framework integrates principled $p$/$e$-value construction, randomized decision-making, and active sampling strategies. Extensive evaluations—including synthetic simulations, genome-wide association studies (GWAS), and large language model–driven clinical prediction tasks—demonstrate substantial gains in statistical power under fixed computational budgets, while preserving scalability and inferential reliability.

Technology Category

Application Category

📝 Abstract
In large-scale hypothesis testing, computing exact $p$-values or $e$-values is often resource-intensive, creating a need for budget-aware inferential methods. We propose a general framework for active hypothesis testing that leverages inexpensive auxiliary statistics to allocate a global computational budget. For each hypothesis, our data-adaptive procedure probabilistically decides whether to compute the exact test statistic or a transformed proxy, guaranteeing a valid $p$-value or $e$-value while satisfying the budget constraint in expectation. Theoretical guarantees are established for our constructions, showing that the procedure achieves optimality for $e$-values and for $p$-values under independence, and admissibility for $p$-values under general dependence. Empirical results from simulations and two real-world applications, including a large-scale genome-wide association study (GWAS) and a clinical prediction task leveraging large language models (LLM), demonstrate that our framework improves statistical efficiency under fixed resource limits.
Problem

Research questions and friction points this paper is trying to address.

Develops budget-aware methods for large-scale hypothesis testing.
Uses auxiliary statistics to allocate computational resources efficiently.
Ensures valid p-values or e-values under budget constraints.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Active hypothesis testing with budget-aware computational allocation
Data-adaptive procedure using inexpensive auxiliary statistics
Guarantees valid p-values or e-values under budget constraints
🔎 Similar Papers
No similar papers found.
Q
Qi Kuang
Department of Statistics and Data Science, Fudan University
B
Bowen Gang
Department of Statistics and Data Science, Fudan University
Y
Yin Xia
Department of Statistics and Data Science, Fudan University