Score
Design and implement resampling schemes that draw contiguous or overlapping blocks of observations to produce bootstrap replicates which preserve serial or batch dependence and autocorrelation. Use those replicates to build unbiased or bias‑reduced bootstrap estimators and to compute confidence intervals, p‑values, and uncertainty quantification for statistics computed on dependent batches—including success‑rate metrics analogous to pass@k—and to support hypothesis testing and fair evaluation of correlated samplers.
This study addresses the unreliable estimation of repeatability, between-laboratory, and reproducibility variance components under ISO 5725 standards when sample sizes are small or variance structures are extreme. To overcome this limitation, the authors propose a tailored Bootstrap resampling strategy adapted to a one-way random effects model. The approach refines point estimates by adjusting within-laboratory resampling and constructs confidence intervals via a two-stage resampling scheme integrated with bias-corrected and accelerated (BCa) techniques. Extensive simulations and validation using real data from ISO 5725-4 demonstrate that the proposed method substantially improves estimation accuracy and confidence interval coverage. It yields reliable, near-nominal or conservatively valid inferences for small- to moderate-sized experiments and clearly delineates optimal strategies across different practical scenarios.
This paper addresses the lack of generality and theoretical foundations in existing bootstrap hypothesis testing frameworks. We propose a unified bootstrap testing framework that accommodates both null-distribution-based resampling and diverse nonstandard bootstrap schemes. We first systematically characterize the exchangeability condition and statistical functional construction criteria, prove the local asymptotic equivalence of different resampling schemes in terms of statistical power, and identify the intrinsic mechanism behind the failure of the naive bootstrap. Leveraging empirical process theory and weak convergence analysis, we rigorously establish the asymptotic exactness and consistency of the test under fixed alternatives. An accompanying open-source R package, *BootstrapTests*, validates the theoretical properties in independence testing, linear regression coefficient testing, and copula model goodness-of-fit testing. Finite-sample simulations demonstrate that the proposed method significantly improves statistical power.
Bootstrapping on massive datasets faces bottlenecks including high computational cost, substantial distributed communication overhead, and memory constraints. To address these challenges, this paper proposes an MPI-based scalable parallel bootstrapping algorithm. Its key contributions are: (1) a local sufficient statistic aggregation mechanism that drastically reduces inter-node communication volume; (2) a synchronized pseudorandom number generation strategy that eliminates the need for storing massive intermediate resampling data, thereby alleviating memory pressure; and (3) rigorous theoretical modeling of communication and computational complexity to ensure scalability. Experimental results demonstrate that the method maintains statistical consistency while significantly reducing both communication overhead and memory footprint. It substantially outperforms naive parallel baselines and is well-suited for ultra-large-scale data analytics.
This paper addresses the complexity and high pedagogical/practical barriers associated with conventional uncertainty quantification methods—such as standard errors, confidence intervals, and hypothesis tests—in statistical inference. To evaluate the potential of nonparametric bootstrap as a unified alternative, we conduct a large-scale simulation study rigorously comparing single bootstrap, double bootstrap, and classical methods across multiple dimensions: sample size, confidence level, data-generating mechanisms, and statistical functionals. Results demonstrate that the double bootstrap consistently achieves superior coverage accuracy, stability, and robustness—particularly under small-sample and non-normal conditions—outperforming both classical approaches and the single bootstrap. We thus establish the double bootstrap as a principled, parsimonious, and high-performance paradigm for uncertainty quantification, providing both theoretical justification and empirical evidence to support its adoption in statistical education and applied practice.
Traditional $n$-out-of-$n$ bootstrap fails for inconsistent estimators—such as extremes, quantiles, and nonsmooth $M$-estimators—due to asymptotic non-normality. To address this, we propose an automated implementation framework for the $m$-out-of-$n$ bootstrap. Our key methodological contribution is the first systematic development of adaptive estimation procedures for both the scaling factor $ au_n$ and the optimal subsample size $m$, grounded in asymptotic theory for inconsistent estimation. We rigorously evaluate multiple $m$-selection strategies via extensive Monte Carlo simulations, assessing their finite-sample coverage accuracy for confidence intervals. Based on this framework, we develop the R package `moonboot`, enabling robust confidence interval construction for diverse inconsistent estimators. Empirical results demonstrate that our approach substantially improves actual coverage probability in small samples, while maintaining theoretical validity and practical usability.
Traditional bootstrap and conformal prediction methods fail in time series settings due to violations of exchangeability and the absence of a unified framework that supports dependence-aware resampling and adaptive conformal calibration. This work proposes the first typed API integrating block, residual, sieve, and wild resampling schemes with adaptive conformal approaches such as EnbPI and ACI, enabling distribution-free uncertainty quantification. Leveraging compilation-based acceleration and streaming reductions, the method requires only O(B) additional memory, circumventing the O(Bn) tensor duplication typical of conventional implementations. Empirical results demonstrate that the approach substantially mitigates undercoverage under the i.i.d. assumption, with sieve resampling achieving coverage closest to the nominal level for short-memory linear processes, while running several times faster than the arch benchmark.
This study addresses the inconsistency and lack of uniform validity in inference for two-way clustered linear regression, which arise from heterogeneous score components and infeasible asymptotic regimes. To resolve these issues, the paper develops a unified, feasible inference framework encompassing four distinct asymptotic mechanisms. It innovatively integrates mechanism adaptivity with flexible modeling of spatial dependence along the first dimension and serial dependence along the second. For the first time in the two-way clustering literature, a data-driven mechanism classifier and a projection-based wild bootstrap procedure are introduced. Theoretical analysis establishes four impossibility results concerning consistency and distinguishability, while Monte Carlo simulations demonstrate that the proposed method achieves both high precision and strong adaptability under complex clustering structures, thereby enabling uniformly valid inference across asymptotic regimes.
Existing resampling-based simultaneous confidence bands for the cumulative hazard function often suffer from undercoverage in finite samples with right-censoring. This work proposes a novel approach that integrates an exchangeable bootstrap with box calibration, preserving the ratio structure of the Nelson–Aalen estimator and constructing upper and lower step-function envelopes over a linearly scanned grid of event times. By doing so, it effectively calibrates vertical deviations without requiring variance-stabilizing transformations. The method substantially improves coverage accuracy in finite samples, achieving empirical coverage probabilities much closer to the nominal level across diverse hazard function shapes and censoring rates. Its practical advantage is further demonstrated through an application to melanoma data, where it yields visibly refined simultaneous confidence bands for the cumulative hazard.
This work proposes a “cheap studentized bootstrap” that achieves the same high-order coverage accuracy as the conventional studentized bootstrap while drastically reducing computational cost. Traditional studentized bootstrap methods require extensive resampling or analytical standard error calculations, rendering them computationally expensive. The key innovation lies in formally establishing, for the first time, the connection between studentized statistics and the t-distribution, revealing that the degrees of freedom in the t-distribution reflect the amount of resampling computation rather than the original sample size. Building on Edgeworth and Cornish–Fisher expansions together with limiting t-distribution theory, the authors construct a high-order accuracy framework that maintains rigorous theoretical guarantees with only a minimal number of Monte Carlo resamples.