🤖 AI Summary
This work addresses the limitations of traditional Monte Carlo tests, which rely on joint exchangeability—a condition often violated in practical settings such as MCMC initialization and posterior predictive checks, where simulated samples typically satisfy only marginal null distributions or pairwise exchangeability, potentially invalidating empirical p-values. Without assuming joint exchangeability or mixing conditions, the paper establishes finite-sample, non-asymptotic validity guarantees by leveraging a conditionally i.i.d. structure, and proves for the first time that the empirical p-value satisfies \( \mathbb{P}(p_m \leq \alpha) \leq 2\alpha \). This result provides a unified theoretical explanation for the inferential behavior of various computational testing procedures, revealing that they are at most twice the nominal significance level in finite samples, thereby offering rigorous justification for their practical use.
📝 Abstract
In hypothesis testing, Monte Carlo tests are usually justified either by exact null simulation or by joint exchangeability of the observed data and its simulated copies. This leaves a gap for common computational procedures, such as parallel MCMC sampling initialized at the observed data, where each copy may be marginally null and even pairwise exchangeable with the observation, but the full collection is not jointly exchangeable. In such cases the usual empirical p-value can be invalid when the chain has not mixed, while exactly exchangeable constructions such as the Besag--Clifford hub-and-spoke sampler may suffer from high conditional Monte Carlo variability. We give finite-sample guarantees for this intermediate regime. If, under the null, the observed data $X$ and a copy $X'\sim P(\cdot\mid X)$ are conditionally i.i.d.\ given a latent variable, then for any prespecified statistic and any finite number $m$ of conditionally independent Monte Carlo copies, the resulting empirical p-value obeys $\mathbb P\{p_m\le α\}\le 2α$. This guarantee requires no mixing conditions and holds for any number of copies $m$, and it explains finite-sample oscillatory behavior in inference via MCMC sampling. In addition, we further show that the guarantee provides insights into inference problems arising in other settings, including inference on Bayesian models (recovering a classical result showing validity up to a factor of $2$ for posterior predictive p-values), and inference via balanced permutation tests.