Score
Designs, builds, and analyzes statistical decision procedures and calibration rules that bound the probability of Type I (false positive) errors in hypothesis testing, including selection of significance thresholds, stopping rules, and multiple-comparison adjustments. Works on tradeoffs between Type I control and other objectives such as Type II error rates and sampling cost, and on methods that guarantee specified false-positive guarantees (per-test, family-wise, or false-discovery metrics).
In small-sample two-sample binomial proportion testing, conventional methods suffer from a power imbalance: the Wald test is overly liberal (inflated Type I error), while Fisher’s exact test is overly conservative (reduced power). Method: This paper proposes a novel exact two-sample testing framework based on integer programming. Extending a classic 1969 approach, it operates directly on the discrete null sample space to strictly control Type I error, and constructs the rejection region via combinatorial optimization to maximize weighted average power—supporting customizable prior weights. Contribution/Results: We prove that the method guarantees exact Type I error control and achieves average power that is optimal or near-optimal under small samples. Empirical evaluations demonstrate substantial improvements over standard benchmarks—including Wald and Fisher’s exact tests—yielding superior balance between statistical power and robustness.
This paper studies strategic hypothesis testing within a principal–agent framework: the agent holds private beliefs about product efficacy and may manipulate submitted data to maximize expected payoff; the principal must design a p-value threshold to balance Type I and Type II error risks. Methodologically, it innovatively integrates game-theoretic reasoning with classical statistical hypothesis testing by imposing incentive-compatibility constraints. The analysis establishes that the optimal p-value threshold exhibits a monotonic, analytically tractable structure in the agent’s strategic behavior, yielding a closed-form solution. Theoretically, it demonstrates that regulators can endogenously mitigate strategic reporting by calibrating the critical p-value, thereby unifying statistical robustness with incentive compatibility. Empirical validation using FDA drug approval data confirms the model’s predictive power, providing regulators with an interpretable and computationally tractable framework for optimizing approval policies.
In Bayesian sequential trials, error rate evaluation relies on computationally expensive Monte Carlo simulations, hindering efficient optimization of sample size and decision thresholds. Method: This paper establishes, for the first time, analytical functional relationships between posterior and posterior predictive probabilities and sample size. Leveraging Bayesian decision theory and asymptotic analysis—combined with numerical fitting and error-rate inversion—the method enables precise error-rate assessment for any sample size using only two simulations, and rapidly identifies optimal design parameters. Contribution/Results: The approach drastically reduces computational cost while achieving error-rate control accuracy comparable to conventional simulation-based methods. In two real-world case studies, it attains exact error-rate calibration and accelerates design optimization by several orders of magnitude. This provides a scalable, verifiable, and highly efficient design paradigm for Bayesian adaptive trials.
Classical false discovery rate (FDR) control methods, such as Benjamini–Hochberg (BH), rely on stringent pointwise control of Type I error (strong control), limiting their applicability under weaker inferential assumptions. This work addresses FDR control when only average-level (i.e., weak) control of significance level is required across tests. Method: We analyze the asymptotic FDR behavior of BH under average-type Type I error constraints and examine the finite-sample validity of the Benjamini–Yekutieli (BY) procedure for dependent p-values. Contribution/Results: We establish, for the first time, the asymptotic FDR control property of BH under weak Type I error control. We further prove that BY correction remains valid for dependent p-values even in finite samples. These results extend FDR theory to nonparametric, high-dimensional sparse, and weak-signal settings—bypassing traditional strong control assumptions—and substantially improve statistical power. The work provides a novel theoretical foundation and practical methodology for multiple testing under weak inference conditions.
This paper identifies a previously unrecognized inflation of Type I error rate under conventional optimal allocation rules in response-adaptive clinical trials with binary outcomes. To address this, we propose two novel optimal allocation schemes—based on the score test and finite-sample parameter estimation—that jointly optimize statistical power and patient benefit (i.e., minimize expected treatment failures). Our approach avoids Wald-type tests relying on unknown true parameters, thereby enhancing small-sample robustness. Monte Carlo simulations demonstrate that the proposed methods strictly control Type I error while significantly improving patient outcomes—reducing failure rates—in both early-phase and confirmatory trials. The framework naturally extends to multi-arm designs and continuous outcomes, providing both theoretical foundations and practical tools for response-adaptive trial design.
This work addresses the trade-off between confirmation bias and high uncertainty rates in sequential hypothesis testing, where conventional methods either couple stopping rules with decision criteria—introducing bias—or fully decouple them, yielding excessive inconclusive outcomes. To reconcile reliability and decisiveness, the authors propose DPitG, a novel Bayesian approach that jointly incorporates a precision target (posterior highest density interval [HDI] width ≤ ω) and a definitive decision rule (HDI entirely within or outside the region of practical equivalence [ROPE]) into its stopping criterion. Built upon the HDI-ROPE framework, DPitG accommodates both binary and continuous data and provides a closed-form sample size planning formula. In simulations of a fair coin test, DPitG reduces the rate of uncertain conclusions from 62% to 2% with only a 5% increase in median sample size, achieves zero false positives, and maintains conclusion certainty above 97%, substantially outperforming existing methods.
This study investigates the admissibility and complete class problems for false discovery rate (FDR) control procedures within the e-value framework. Drawing on statistical decision theory, it introduces strong and weak dominance relations to establish, for the first time, a theoretical foundation for admissibility in e-value-based multiple testing with FDR control. The main contributions include proving that every step-down procedure is strongly dominated by some weighted average eBH procedure; demonstrating that weighted average eBH procedures without constant terms are admissible at any FDR level; and showing that, under symmetry, this class of procedures forms a complete class, with its members being maximal only when the FDR threshold is sufficiently small—thereby establishing their structural optimality.
This study addresses the lack of a systematic taxonomy in Bayesian A/B testing, which has led to the conflation of prior selection and stopping rules, resulting in methodological misuse and performance risks. The authors propose a three-tier classification framework encompassing posterior consistency, error rate control under Bayes factor–based stopping, and empirical Bayes–driven false discovery rate calibration. This work provides the first comprehensive formalization of Bayesian A/B testing methodologies, demonstrating that Bayes factor stopping is approximately optimal across a range of loss functions and establishing empirical Bayes as the sole viable route to achieving third-tier calibration. Simulations reveal that flat priors combined with posterior-based stopping amount to unprincipled peeking, that well-calibrated empirical Bayes priors substantially reduce estimation error, and that expected loss–based stopping minimizes regret only when the deployment cost of null effects is negligible.
This study addresses the critical issue that existing selective prediction methods in signal domains—such as anomalous sound detection and AI-generated image forensics—often yield a false sense of security due to the use of uncalibrated thresholds, resulting in actual error rates that substantially exceed users’ prescribed risk budgets. The work presents the first systematic audit of four distribution-free calibration rules (NAIVE, Hoeffding, Clopper–Pearson, and Betting) regarding their risk control performance on both real and synthetic data. Findings reveal that NAIVE exceeds the risk budget in 49–73% of experiments; Clopper–Pearson and Betting achieve zero violations under exchangeability but suffer 9–30% violation rates when deployed in grouped settings where exchangeability fails. Group-wise thresholding restores valid risk control at the cost of reduced coverage. The study underscores the pivotal role of tight confidence bounds for effective coverage and identifies uncalibrated thresholds as the root cause of risk miscontrol.