Score
Statistical methods for designing hypothesis tests and estimators that rigorously control the type I (false positive) error rate at a preset level while optimizing power and sampling cost and assessing finite‑sample and asymptotic properties.
In multiple hypothesis testing, maximizing statistical power under strict control of the family-wise error rate (FWER) has long been hindered by the absence of efficient, tractable algorithms for solving the associated dual optimization problem. This paper derives, for the first time, the necessary optimality conditions for the FWER-constrained optimal testing dual problem. Building on Lagrangian duality and convex analysis, we propose the first coordinate descent algorithm with guaranteed linear convergence—achieving complexity O(log(1/ε))—thereby filling a critical computational gap. Our method unifies statistical power maximization with computational efficiency. Extensive simulations demonstrate substantially higher power than state-of-the-art alternatives. Furthermore, applications to real clinical and financial datasets successfully identify novel, statistically significant signals not detected by existing approaches.
Balancing statistical power and computational efficiency remains challenging in multiple testing, particularly under stringent family-wise error rate (FWER) control. Method: We propose a bottom-up construction framework grounded in closed testing, the first to embed a global power-optimization objective directly into the design of local tests at each hierarchical level. This yields a novel testing procedure that satisfies consonance and strictly controls FWER. By recursively constructing analytically tractable and computationally feasible test statistics, our method ensures statistical rigor while markedly improving computational scalability. Contribution/Results: Extensive simulations and real-data analyses demonstrate that our procedure achieves significantly higher true positive rates than leading competitors—including Holm, Hochberg, and existing closed testing variants—while maintaining exact FWER control. Theoretical analysis confirms its optimality properties, and empirical results underscore its practical utility across diverse applications.
This work addresses the finite-sample minimax robust hypothesis testing problem under distributional uncertainty by rigorously bridging finite-sample and asymptotic optimal solutions through asymptotic theory, thereby avoiding conventional heuristic constructions. The authors model uncertainty using total variation distance and band models, and derive explicit parametric forms for the least favorable distributions and the robust likelihood ratio function. Theoretical analysis establishes that, under both uncertainty models, the finite-sample minimax robust test coincides with its asymptotic counterpart, extending existing results to asymmetric robust parameter settings and providing a systematic unification of prior approaches. Numerical simulations corroborate the theoretical guarantees and demonstrate the practical efficacy of the proposed test.
In nonstandard testing scenarios, evaluating the optimality of heuristic test statistics is challenging. Method: This paper proposes a nested numerical optimization framework to assess whether a given heuristic test is approximately optimal—i.e., whether its power curve approximates the power envelope generated by a weighted average power (WAP)-maximizing test. Contribution/Results: We introduce, for the first time, a data-driven approach to approximate the optimal weighting function. Crucially, we establish theoretically that the rejection probability of the WAP-optimal test itself constitutes a tight upper bound on the power of any heuristic test—a result both theoretically unexpected and practically valuable. The method is provably convergent and is successfully applied to two canonical problems: robust conditional likelihood ratio (CLR) testing under weak instruments and testing at nuisance parameter boundaries. Empirical results confirm the near-optimality of the heuristic tests in these settings.
This paper identifies a previously unrecognized inflation of Type I error rate under conventional optimal allocation rules in response-adaptive clinical trials with binary outcomes. To address this, we propose two novel optimal allocation schemes—based on the score test and finite-sample parameter estimation—that jointly optimize statistical power and patient benefit (i.e., minimize expected treatment failures). Our approach avoids Wald-type tests relying on unknown true parameters, thereby enhancing small-sample robustness. Monte Carlo simulations demonstrate that the proposed methods strictly control Type I error while significantly improving patient outcomes—reducing failure rates—in both early-phase and confirmatory trials. The framework naturally extends to multi-arm designs and continuous outcomes, providing both theoretical foundations and practical tools for response-adaptive trial design.
This study addresses the lack of a statistically powerful procedure for block-structured multiple hypothesis testing under strong family-wise error rate (FWER) control. The authors propose the first framework that achieves optimal statistical power while guaranteeing strong FWER control, applicable to settings such as gatekeeping trials, dose-finding studies, and multi-tissue eQTL mapping. The method delivers the first provably power-optimal solution for blocks of size three, integrating a KKT-condition-driven marginal power balancing algorithm, closure principle acceleration, sample-splitting plug-in estimation, and a Sidák-type refinement to ensure finite-sample validity and robustness to heterogeneous block structures and unknown alternative distributions. Empirical evaluations demonstrate 1.4–1.7× higher power than the strongest existing baselines across diverse correlation and sparsity regimes, and in real-world eQTL and A/B testing data, it identifies an order-of-magnitude more significant full-block discoveries.
Traditional hypothesis testing relies on fixed sample sizes and struggles to accommodate sequentially arriving data, while existing sequential methods either require rigid pre-specified analysis plans or compromise statistical power. This work proposes a general framework that transforms any fixed-sample test into a sequentially valid test usable at arbitrary stopping times. By modeling future observations as missing data and predicting, under the null hypothesis, the probability that the full-sample test would reject, the method constructs an adaptive stopping rule. It requires no pre-specified analysis plan, rigorously controls the Type I error rate, and achieves near-optimal statistical power. Moreover, under the alternative hypothesis, it substantially reduces the required sample size, making it especially suitable for applications such as clinical trials where early stopping for efficacy or futility is critical.
This work addresses the problem of multiple hypothesis testing for edge distributions across multiple data streams. It proposes a sequential testing procedure that, for the first time, systematically incorporates arbitrary forms of prior information about the configuration of true and false hypotheses—such as known values or lower bounds on the number of active streams under each hypothesis, or mutual exclusivity constraints—while rigorously controlling the familywise error rate. By integrating sequential analysis with a search strategy over minimal alternative hypothesis configurations, the method achieves asymptotic optimality in terms of expected sample size among all valid procedures, without compromising reliability. Theoretical analysis establishes its computational efficiency and asymptotic optimality, and numerical experiments further demonstrate its substantial advantages in both testing efficiency and accuracy.
This paper investigates the fundamental reason why randomization tests fail to achieve exact Type I error control for common null hypotheses—such as zero-mean—in finite samples. Method: Drawing on group action theory, the authors develop a unified analytical framework and derive necessary and sufficient conditions for a null hypothesis to admit exact finite-sample randomization tests. Contribution/Results: They prove that exact finite-sample randomization testing is possible if and only if the null hypothesis is distributionally invariant under some nontrivial linear group action. Consequently, asymmetric hypotheses—including the zero-mean null—are rigorously shown to be inherently infeasible for exact randomization testing. This result clarifies the intrinsic limitations of randomization inference and underscores the necessity of relying on asymptotic methods where exact finite-sample control is unattainable. The work thus establishes precise theoretical boundaries for valid application of randomization-based inference.
This work addresses the challenge of controlling the false discovery rate (FDR) in large-scale multiple hypothesis testing when only a limited number of null samples are available and hypotheses exhibit arbitrary structural dependencies. The authors propose a unified framework that leverages reproducing kernel Hilbert spaces (RKHS) to model the intrinsic structure among hypotheses, integrates uncertainty quantification of finite-sample p-values, and extends mirror statistics to the counting space. Within this framework, they construct two decision rules that provide rigorous FDR guarantees. Their approach is the first to simultaneously handle data scarcity and complex dependency structures, achieving substantially improved statistical power while maintaining robustness. Additionally, it offers an efficient strategy for allocating scarce null-distribution samples, enabling a flexible trade-off between precision and power in structured multiple testing.