Score
Designs and implements hypothesis tests and inference procedures based on the random assignment or permutation of units, including Fisher randomization tests and algorithms to compute exact permutation p‑values. Builds and analyzes finite-sample–valid randomization tests that integrate external controls and maintain type I error control under model misspecification.
This paper addresses the low statistical power and poor resolution of conventional placebo tests in synthetic control method (SCM) causal inference under small-sample settings—particularly when α = 0.05 and the number of donor units *N* is small. We propose a leave-two-out randomization inference framework that rigorously controls Type I error rates in finite samples while substantially improving test resolution and statistical power, even under stringent significance levels (α < 1/*N*). Unlike permutation or rank-based tests, our framework accommodates non-uniform treatment assignment and integrates formal sensitivity analysis for robust causal inference. Empirical results demonstrate that, under moderate effect sizes, the proposed method achieves lower actual Type I error rates and higher statistical power compared to standard approaches.
To address the computational bottleneck of performing thousands of hypothesis tests on high-dimensional genetic or neuroimaging data, this paper proposes an anytime-terminating Monte Carlo p-value construction method—the first to extend anytime-valid sequential testing theory to the multiple testing framework. The method is compatible with standard false discovery rate (FDR) control procedures such as Benjamini–Hochberg, guarantees finite-sample FDR control, and substantially reduces the average number of permutations required. Its core innovations integrate sequential Monte Carlo testing, arbitrary stopping time theory, randomized permutation mechanisms, and adaptive correction strategies. Experiments on both synthetic and real-world datasets demonstrate improved statistical power and over 50% reduction in computational time compared to state-of-the-art methods. The implementation is publicly available.
Existing exact confidence interval methods for the average treatment effect (ATE) with binary outcomes in randomized experiments are computationally expensive and often incorrectly assume a binomial distribution. Method: We propose the first exact, asymptotics-free confidence interval construction framework that dispenses with the binomial assumption. Our approach models the finite-population distribution using the hypergeometric distribution and combines combinatorial counting with boundary search, accelerated via a divide-and-conquer optimization strategy. Contribution/Results: The algorithm achieves $O(n log n)$ time complexity, enabling millisecond-scale computation for $n leq 1000$—over 100× faster than prior exact methods—while rigorously guaranteeing nominal coverage. It is especially suited for small-sample settings where asymptotic approximations fail and exact inference is critical.
Traditional linear-model t/F-tests assume fixed sample sizes, rendering them unsuitable for sequential A/B testing requiring continuous monitoring and early stopping—leading to uncontrolled Type-I error inflation. This paper proposes an anytime-valid causal inference framework under linear regression adjustment, introducing the first closed-form anytime-valid F-tests and confidence sequences for both parametric and nonparametric settings. Without imposing strong modeling assumptions, the method guarantees uniform Type-I error control and valid confidence coverage throughout the entire sequential experiment under standard randomized designs. All test statistics are directly computable from standard regression outputs, enabling real-time significance assessment and dynamic confidence interval updating. Deployed on Netflix’s industrial-scale A/B testing platform, the method supports regression-adjusted sequential analysis using pre-treatment covariates, effectively mitigating p-hacking.
In adaptive experiments, data-driven design adjustments invalidate conventional statistical inference; existing methods suffer from narrow applicability and strong assumptions. This paper proposes a selective randomization inference framework: it models the data-generating process via a directed acyclic graph (DAG) and implements conditional-on-selection inference within randomization tests. It is the first systematic integration of this principle into randomization-based inference—requiring neither i.i.d. nor parametric modeling assumptions, and accommodating arbitrary adaptive experimental designs. To address disconnected confidence intervals, we innovatively introduce a holdout-unit method. Theoretically and empirically, our approach strictly controls selective Type-I error and constructs valid confidence intervals for homogeneous treatment effects. It substantially outperforms conventional methods in both robustness and generality.
This work addresses the challenge of finite-sample inference for individual regression coefficients in fixed-design linear models when errors exhibit dependence or heteroskedasticity. The authors propose a unified randomization testing framework based on group permutations, which rigorously controls Type I error under exchangeable errors and enhances power through design-dependent geometric separation. The approach is further extended to non-exchangeable settings, establishing quantitative robustness for approximately symmetric errors. The study proves that the resulting Type I error bound of level $2\alpha$ is tight and, by integrating a constructive algorithm, sub-Gaussian analysis, and conformal inference, achieves substantially improved power under heavy-tailed designs while preserving finite-sample validity.
This study addresses the limitations of permutation testing caused by a finite number of permutations, which often yields empirical p-values of zero or a coarse distribution, thereby compromising statistical validity and undermining multiple testing correction. To overcome this, the authors propose a tail modeling approach based on the Generalized Pareto Distribution (GPD), introducing a support constraint in GPD fitting for the first time to guarantee non-zero and valid extrapolated p-values. The framework integrates robust maximum likelihood estimation, data-driven threshold selection, and a hybrid treatment of discrete–continuous p-values to deliver a complete and reliable approximation scheme. Evaluations on both simulated and real-world single-cell RNA-seq and microbiome datasets demonstrate that the method produces accurate, robust, smooth, and interpretable p-value distributions—even with a limited number of permutations.
This work addresses the challenge that, under non-exchangeability, the covariance structure of permutation statistics deviates from that of the original test statistics, rendering conventional studentization incapable of recovering the correct joint asymptotic distribution. To overcome this limitation, the authors propose a general and computationally efficient covariance correction method that requires no assumptions about specific parameters, test statistics, or permutation schemes, and remains valid even in singular covariance settings. Unlike existing approaches—such as pre-pivoting—which suffer from high computational costs, the proposed method accurately restores the asymptotic dependence structure of permutation statistics. Theoretical analysis and extensive simulations demonstrate that it achieves asymptotically valid and powerful multiple testing across diverse scenarios, significantly outperforming current methods in inferential accuracy and efficiency.
This study addresses the challenge of accurately inferring the distribution of individual treatment effects—such as the proportion benefiting, the median effect, or the maximum impact—in randomized experiments, without suffering power loss due to suboptimal pre-specified test statistics. The authors propose an adaptive randomization test that combines multiple rank-based statistics, ensuring finite-sample validity without requiring prior knowledge of the optimal statistic. Innovatively integrating adaptive statistic combination with stratified weighting, the method effectively circumvents the power degradation typically induced by multiple comparison corrections and accommodates heterogeneous stratified experimental designs. In an empirical application to a teacher training program, the approach reveals that approximately half of the teachers experience significant benefits, demonstrating superior detection power and interpretability compared to conventional single rank-based tests.
This study addresses the computational burden of exact randomization tests for constructing confidence intervals of the average treatment effect in randomized experiments with binary outcomes. The authors propose an efficient algorithm that achieves exact inference under balanced Bernoulli and matched-pair designs using only $O(\log n)$ randomization tests, yielding an exponential speedup over brute-force approaches. They further establish that this complexity is information-theoretically optimal. The work also uncovers fundamental differences in computational complexity across randomization schemes: complete randomization requires $O(n \log n)$ tests, while general Bernoulli designs necessitate $O(n^2)$. This paper presents the first systematic theory of computational complexity for randomization-based inference and provides practical algorithms that match the derived lower bounds.