type i error control

Statistical methods for designing hypothesis tests and estimators that rigorously control the type I (false positive) error rate at a preset level while optimizing power and sampling cost and assessing finite‑sample and asymptotic properties.

typeierrorcontrol

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

In multiple hypothesis testing, maximizing statistical power under strict control of the family-wise error rate (FWER) has long been hindered by the absence of efficient, tractable algorithms for solving the associated dual optimization problem. This paper derives, for the first time, the necessary optimality conditions for the FWER-constrained optimal testing dual problem. Building on Lagrangian duality and convex analysis, we propose the first coordinate descent algorithm with guaranteed linear convergence—achieving complexity O(log(1/ε))—thereby filling a critical computational gap. Our method unifies statistical power maximization with computational efficiency. Extensive simulations demonstrate substantially higher power than state-of-the-art alternatives. Furthermore, applications to real clinical and financial datasets successfully identify novel, statistically significant signals not detected by existing approaches.

Demonstrates linear convergence and superior power in simulations and real-world applicationsDerives necessary optimality conditions for dual optimization in multiple hypothesis testingProposes an efficient coordinate-wise algorithm for computing the optimal dual solution

The Bottom-Up Approach for Powerful Testing with FWER Control

Nov 17, 2025
RK
Rajesh Karmakar
🏛️ Tel-Aviv University

Balancing statistical power and computational efficiency remains challenging in multiple testing, particularly under stringent family-wise error rate (FWER) control. Method: We propose a bottom-up construction framework grounded in closed testing, the first to embed a global power-optimization objective directly into the design of local tests at each hierarchical level. This yields a novel testing procedure that satisfies consonance and strictly controls FWER. By recursively constructing analytically tractable and computationally feasible test statistics, our method ensures statistical rigor while markedly improving computational scalability. Contribution/Results: Extensive simulations and real-data analyses demonstrate that our procedure achieves significantly higher true positive rates than leading competitors—including Holm, Hochberg, and existing closed testing variants—while maintaining exact FWER control. Theoretical analysis confirms its optimality properties, and empirical results underscore its practical utility across diverse applications.

Design novel multiple testing procedures with strong FWER controlOptimize power objectives for true individual discoveries efficientlyPropose bottom-up approach for computationally practical improved power

This work addresses the finite-sample minimax robust hypothesis testing problem under distributional uncertainty by rigorously bridging finite-sample and asymptotic optimal solutions through asymptotic theory, thereby avoiding conventional heuristic constructions. The authors model uncertainty using total variation distance and band models, and derive explicit parametric forms for the least favorable distributions and the robust likelihood ratio function. Theoretical analysis establishes that, under both uncertainty models, the finite-sample minimax robust test coincides with its asymptotic counterpart, extending existing results to asymmetric robust parameter settings and providing a systematic unification of prior approaches. Numerical simulations corroborate the theoretical guarantees and demonstrate the practical efficacy of the proposed test.

asymptotic theorydistributional uncertaintyfinite-sample

Numerical Analysis of Test Optimality

Dec 22, 2025
PK
Philipp Ketz
🏛️ Paris School of Economics | University of Colorado | University of Bonn

In nonstandard testing scenarios, evaluating the optimality of heuristic test statistics is challenging. Method: This paper proposes a nested numerical optimization framework to assess whether a given heuristic test is approximately optimal—i.e., whether its power curve approximates the power envelope generated by a weighted average power (WAP)-maximizing test. Contribution/Results: We introduce, for the first time, a data-driven approach to approximate the optimal weighting function. Crucially, we establish theoretically that the rejection probability of the WAP-optimal test itself constitutes a tight upper bound on the power of any heuristic test—a result both theoretically unexpected and practically valuable. The method is provably convergent and is successfully applied to two canonical problems: robust conditional likelihood ratio (CLR) testing under weak instruments and testing at nuisance parameter boundaries. Empirical results confirm the near-optimality of the heuristic tests in these settings.

Applies method to weak instrument and boundary nuisance parameter testsDevelops a numerical framework to assess ad hoc test optimalityUses nested optimization to approximate weight functions for power comparison

Revisiting Optimal Allocations for Binary Responses: Insights from Considering Type-I Error Rate Control

Feb 10, 2025
LP
Lukas Pin
🏛️ University of Cambridge | George Mason University

This paper identifies a previously unrecognized inflation of Type I error rate under conventional optimal allocation rules in response-adaptive clinical trials with binary outcomes. To address this, we propose two novel optimal allocation schemes—based on the score test and finite-sample parameter estimation—that jointly optimize statistical power and patient benefit (i.e., minimize expected treatment failures). Our approach avoids Wald-type tests relying on unknown true parameters, thereby enhancing small-sample robustness. Monte Carlo simulations demonstrate that the proposed methods strictly control Type I error while significantly improving patient outcomes—reducing failure rates—in both early-phase and confirmatory trials. The framework naturally extends to multi-arm designs and continuous outcomes, providing both theoretical foundations and practical tools for response-adaptive trial design.

Explores and critiques existing approaches for controlling type-I error ratesInvestigates type-I error rate inflation in optimal response-adaptive designsProposes two new optimal allocation methods for binary outcomes

Latest Papers

What's happening recently
View more

This study addresses the lack of a statistically powerful procedure for block-structured multiple hypothesis testing under strong family-wise error rate (FWER) control. The authors propose the first framework that achieves optimal statistical power while guaranteeing strong FWER control, applicable to settings such as gatekeeping trials, dose-finding studies, and multi-tissue eQTL mapping. The method delivers the first provably power-optimal solution for blocks of size three, integrating a KKT-condition-driven marginal power balancing algorithm, closure principle acceleration, sample-splitting plug-in estimation, and a Sidák-type refinement to ensure finite-sample validity and robustness to heterogeneous block structures and unknown alternative distributions. Empirical evaluations demonstrate 1.4–1.7× higher power than the strongest existing baselines across diverse correlation and sparsity regimes, and in real-world eQTL and A/B testing data, it identifies an order-of-magnitude more significant full-block discoveries.

block-structured multiplicityhypothesis testingmultiple testing

Traditional hypothesis testing relies on fixed sample sizes and struggles to accommodate sequentially arriving data, while existing sequential methods either require rigid pre-specified analysis plans or compromise statistical power. This work proposes a general framework that transforms any fixed-sample test into a sequentially valid test usable at arbitrary stopping times. By modeling future observations as missing data and predicting, under the null hypothesis, the probability that the full-sample test would reject, the method constructs an adaptive stopping rule. It requires no pre-specified analysis plan, rigorously controls the Type I error rate, and achieves near-optimal statistical power. Moreover, under the alternative hypothesis, it substantially reduces the required sample size, making it especially suitable for applications such as clinical trials where early stopping for efficacy or futility is critical.

anytime-valid inferencefixed-sample testsequential testing

This work addresses the problem of multiple hypothesis testing for edge distributions across multiple data streams. It proposes a sequential testing procedure that, for the first time, systematically incorporates arbitrary forms of prior information about the configuration of true and false hypotheses—such as known values or lower bounds on the number of active streams under each hypothesis, or mutual exclusivity constraints—while rigorously controlling the familywise error rate. By integrating sequential analysis with a search strategy over minimal alternative hypothesis configurations, the method achieves asymptotic optimality in terms of expected sample size among all valid procedures, without compromising reliability. Theoretical analysis establishes its computational efficiency and asymptotic optimality, and numerical experiments further demonstrate its substantial advantages in both testing efficiency and accuracy.

familywise errorhypothesis configurationmultiple hypotheses

Limitations of Randomization Tests in Finite Samples

Dec 07, 2025
DD
Deniz Dutz
🏛️ University of Chicago | Columbia University

This paper investigates the fundamental reason why randomization tests fail to achieve exact Type I error control for common null hypotheses—such as zero-mean—in finite samples. Method: Drawing on group action theory, the authors develop a unified analytical framework and derive necessary and sufficient conditions for a null hypothesis to admit exact finite-sample randomization tests. Contribution/Results: They prove that exact finite-sample randomization testing is possible if and only if the null hypothesis is distributionally invariant under some nontrivial linear group action. Consequently, asymmetric hypotheses—including the zero-mean null—are rigorously shown to be inherently infeasible for exact randomization testing. This result clarifies the intrinsic limitations of randomization inference and underscores the necessity of relying on asymptotic methods where exact finite-sample control is unattainable. The work thus establishes precise theoretical boundaries for valid application of randomization-based inference.

Identifies when null hypotheses allow valid randomization testsProves certain nulls like mean zero lack valid randomization testsShows linear group tests only apply to symmetric or normal distributions

This work addresses the challenge of controlling the false discovery rate (FDR) in large-scale multiple hypothesis testing when only a limited number of null samples are available and hypotheses exhibit arbitrary structural dependencies. The authors propose a unified framework that leverages reproducing kernel Hilbert spaces (RKHS) to model the intrinsic structure among hypotheses, integrates uncertainty quantification of finite-sample p-values, and extends mirror statistics to the counting space. Within this framework, they construct two decision rules that provide rigorous FDR guarantees. Their approach is the first to simultaneously handle data scarcity and complex dependency structures, achieving substantially improved statistical power while maintaining robustness. Additionally, it offers an efficient strategy for allocating scarce null-distribution samples, enabling a flexible trade-off between precision and power in structured multiple testing.

False Discovery RateFinite Null SamplesMultiple Hypothesis Testing

Hot Scholars