Score
Designs and implements permutation-based hypothesis tests and associated statistics — including paired and two-sample tests, pairwise and nodewise comparisons, joint-statistic and split testing — by constructing appropriate label-exchange or sample-permutation schemes and computing permutation distributions and empirical p‑values. Controls type I error and multiple comparisons (for example via Bonferroni-corrected +1 Monte Carlo p‑values), and applies practical computation strategies such as adaptive stopping, limited threshold searches, and other permutation test design choices to manage runtime while ensuring valid inference.
To address the computational bottleneck of performing thousands of hypothesis tests on high-dimensional genetic or neuroimaging data, this paper proposes an anytime-terminating Monte Carlo p-value construction method—the first to extend anytime-valid sequential testing theory to the multiple testing framework. The method is compatible with standard false discovery rate (FDR) control procedures such as Benjamini–Hochberg, guarantees finite-sample FDR control, and substantially reduces the average number of permutations required. Its core innovations integrate sequential Monte Carlo testing, arbitrary stopping time theory, randomized permutation mechanisms, and adaptive correction strategies. Experiments on both synthetic and real-world datasets demonstrate improved statistical power and over 50% reduction in computational time compared to state-of-the-art methods. The implementation is publicly available.
This paper addresses the often-overlooked exchangeability assumption underlying permutation tests in multiple linear regression—a condition critical for valid statistical inference. We rigorously clarify the logical relationship between exchangeability and the null hypothesis, and systematically evaluate the robustness of common permutation schemes—response permutation, residual permutation, and design-matrix permutation—under both satisfied and violated exchangeability conditions, via theoretical analysis and simulation studies. We propose a novel, pedagogically integrated framework that unifies conceptual understanding with formal theory, and extend the analysis for the first time to hierarchical and clustered regression models, enhancing methodological generality. Results show that while standard regression settings yield consistent conclusions across permutation schemes, inference deteriorates markedly when exchangeability fails. Our framework significantly improves students’ conceptual grasp of resampling-based inference, offering a new paradigm for statistics education and applied practice.
Traditional Monte Carlo methods lack sufficient precision for estimating extremely small p-values in two-sample permutation tests, posing challenges for multiple hypothesis testing correction. This work proposes HAMSTest (Hash-Augmented Multilevel Splitting Test), an adaptive multilevel splitting Monte Carlo algorithm enhanced with hashing techniques. By integrating hash-based state representation with an adaptive splitting strategy, HAMSTest effectively handles the discreteness of test statistic distributions and enables efficient, high-precision estimation of arbitrarily small p-values. The method is applicable to widely used nonparametric tests such as Kolmogorov–Smirnov and Mann–Whitney U, and also supports user-defined test statistics. HAMSTest substantially improves estimation accuracy while maintaining computational efficiency and is made accessible through an open-source Python package, hamstest.
This work addresses the challenge that, under non-exchangeability, the covariance structure of permutation statistics deviates from that of the original test statistics, rendering conventional studentization incapable of recovering the correct joint asymptotic distribution. To overcome this limitation, the authors propose a general and computationally efficient covariance correction method that requires no assumptions about specific parameters, test statistics, or permutation schemes, and remains valid even in singular covariance settings. Unlike existing approaches—such as pre-pivoting—which suffer from high computational costs, the proposed method accurately restores the asymptotic dependence structure of permutation statistics. Theoretical analysis and extensive simulations demonstrate that it achieves asymptotically valid and powerful multiple testing across diverse scenarios, significantly outperforming current methods in inferential accuracy and efficiency.
This paper addresses the problem of testing exchangeability of random variables and invariance under compact groups. Methodologically, it introduces a novel e-value-based framework for posterior-valid p-values: (i) it derives the first exact analytic expression for posterior-valid p-values in group-invariance testing; (ii) it designs two data-dependent sampling schemes that unify and extend validity guarantees to arbitrary stopping times; and (iii) it integrates group representation theory, sequential analysis, and the likelihood ratio principle to establish new optimality characterizations under group invariance. Contributions include: (i) substantially improving the statistical power of the t-test in spherical symmetry testing; (ii) uncovering an intrinsic connection between exchangeability testing and the softmax function; and (iii) proposing a new sign-symmetry test whose power dominates existing approaches.
This study addresses the problem of testing conditional independence between random variables \( X \) and \( Y \) given a confounding variable \( Z \). It proposes a local permutation test based on data-adaptive binning—such as equal-count binning—where permutations of \( X \) and \( Y \) are performed within each subregion defined by \( Z \). The method provides, for the first time, finite-sample Type I error control guarantees for arbitrary test statistics. Under linear confounding models, it achieves power comparable to that of the oracle likelihood ratio test. Theoretical analysis shows that a constant bin size suffices to attain performance on par with increasing bin sizes, and numerical experiments confirm the method’s statistical efficiency and practical utility.
This work addresses the problem of multiple hypothesis testing for edge distributions across multiple data streams. It proposes a sequential testing procedure that, for the first time, systematically incorporates arbitrary forms of prior information about the configuration of true and false hypotheses—such as known values or lower bounds on the number of active streams under each hypothesis, or mutual exclusivity constraints—while rigorously controlling the familywise error rate. By integrating sequential analysis with a search strategy over minimal alternative hypothesis configurations, the method achieves asymptotic optimality in terms of expected sample size among all valid procedures, without compromising reliability. Theoretical analysis establishes its computational efficiency and asymptotic optimality, and numerical experiments further demonstrate its substantial advantages in both testing efficiency and accuracy.
本文研究了通过自适应数据收集解决多重测试问题,提出基于e值的后验抽样方法(e-PS),有效控制错误发现率并减少样本需求。
This study challenges the common belief that increasing the number of permutations in Monte Carlo permutation tests invariably enhances statistical power. By integrating discrete distribution theory, probability theory, and Monte Carlo simulations, the authors rigorously demonstrate—for the first time—that under standard settings, statistical power exhibits a non-monotonic, sawtooth-like behavior due to the discreteness of the underlying distribution, leading to infinitely many decreases as the number of permutations grows. These findings reveal a complex relationship between computational budget and test power, offering crucial theoretical corrections for the practical implementation of permutation tests.
This study investigates the conditions under which generalized permutation tests achieve optimal power under non-uniform randomization distributions. Focusing on Monte Carlo permutation tests based on the mean difference statistic, the work introduces individual- and pair-level dispersion measures to quantify deviations from complete randomization. Theoretical analysis shows that when both dispersion measures asymptotically vanish, the permutation distribution converges to a Gaussian limit, yielding stable critical values and attaining Pitman local asymptotic optimality; otherwise, optimality cannot be guaranteed. The key contribution lies in demonstrating that non-uniform randomization schemes can exploit structural heterogeneity in the data to outperform uniform designs, thereby establishing a rigorous connection among dispersion metrics, distributional convergence, and testing power.