Score
Designs and analyzes estimators and inference procedures that use an auxiliary predictive model f(x) as a control variate to improve efficiency, exploit unlabeled covariates to reduce estimation variance, and produce valid confidence intervals and hypothesis tests while accounting for the predictor’s imperfection. These procedures employ bootstrap resampling or related finite-sample calibration techniques (model-assisted or prediction-powered inference) to obtain reliable uncertainty quantification without relying solely on large-sample asymptotics.
This paper addresses the complexity and high pedagogical/practical barriers associated with conventional uncertainty quantification methods—such as standard errors, confidence intervals, and hypothesis tests—in statistical inference. To evaluate the potential of nonparametric bootstrap as a unified alternative, we conduct a large-scale simulation study rigorously comparing single bootstrap, double bootstrap, and classical methods across multiple dimensions: sample size, confidence level, data-generating mechanisms, and statistical functionals. Results demonstrate that the double bootstrap consistently achieves superior coverage accuracy, stability, and robustness—particularly under small-sample and non-normal conditions—outperforming both classical approaches and the single bootstrap. We thus establish the double bootstrap as a principled, parsimonious, and high-performance paradigm for uncertainty quantification, providing both theoretical justification and empirical evidence to support its adoption in statistical education and applied practice.
In experimental causal inference, design uncertainty—arising from the assignment mechanism—is often overlooked under small sample sizes and heterogeneous treatment effects; conventional causal bootstrap methods apply only to completely randomized designs and average treatment effect estimation. This paper addresses this limitation by introducing integer linear programming into the causal bootstrap framework for the first time, enabling computation of the worst-case copula under generalized assignment mechanisms (e.g., conditional unconfoundedness, bounded confounding) to uniformly calibrate design uncertainty. The method accommodates both linear and quadratic treatment effect estimators and is supported by asymptotic theory establishing its validity. Monte Carlo simulations demonstrate that, in small-scale geographic experiments, the proposed approach substantially improves confidence interval coverage and precision while delivering more robust control of Type I error.
This study addresses the challenge of conducting valid Bayesian inference for target parameters—such as causal effects—in the presence of finite-dimensional nuisance parameters. The authors propose a general framework that integrates Bayesian bootstrap with Dirichlet process priors within an estimating equation approach, explicitly accounting for uncertainty in propensity score estimation. The method is robust to model misspecification and remains valid under only single robustness conditions. It extends the "linked Bayesian bootstrap" to nonstandard Bayesian settings, yielding posterior inferences with favorable frequentist properties. Theoretical analysis demonstrates that the resulting posterior distribution exhibits desirable asymptotic behavior, with credible intervals achieving nominal coverage probabilities in large samples.
In hierarchical randomized experiments, small sample sizes or skewed outcomes often lead to overly conservative Neyman variance estimates and failure of normal approximations, thereby undermining the reliability of inference for weighted average treatment effects. To address this, we propose a randomization-based exact variance estimator and introduce two novel causal bootstrap methods—rank-preserving and constant-treatment-effect counterfactual imputation—achieving second-order accuracy improvements; these are further extended to paired-experiment designs. Our approach is grounded in finite-population asymptotic theory and the randomization-inference framework, and is implemented in the open-source R package *CausalBootstrap*. Simulation studies and empirical applications demonstrate that the proposed methods substantially improve confidence interval coverage and statistical power, effectively mitigating inference bias under small-sample and outcome-skewness conditions.
Traditional asymptotic theory (e.g., the Central Limit Theorem) often fails for complex research problems, and prediction-driven inference lacks flexible, assumption-light tools. Method: We propose PPBoot—a novel method that directly embeds standard bootstrap resampling into the Prediction-Powered Inference (PPI) framework, eliminating reliance on asymptotic normality assumptions. PPBoot integrates arbitrary black-box prediction models with bootstrap-based uncertainty quantification and incorporates bias correction to enable valid statistical inference for arbitrary estimators. Contribution/Results: Compared to CLT-dependent PPI(++) methods, PPBoot achieves comparable or superior performance in multivariate regression and causal effect estimation. Crucially, it remains robust in settings where the CLT breaks down—such as small samples and non-i.i.d. data—while maintaining conceptual simplicity and ease of implementation. PPBoot thus substantially broadens the applicability and reliability of prediction-powered inference.
This study addresses Bayesian inference for low-dimensional target parameters in semiparametric models, particularly under the presence of complex nuisance components that may compromise frequentist properties. To this end, we construct posterior distributions by integrating estimating function methods with nonparametric Bayesian techniques—such as Dirichlet processes and Bayesian bootstrap—under conditions weaker than the classical stochastic equicontinuity assumption. We establish asymptotic normality and consistency of the resulting posterior, rigorously identifying the key assumptions required to guarantee desirable frequentist behavior. The theoretical analysis systematically elucidates how relaxing these assumptions affects inferential performance. Extensive simulations corroborate the effectiveness of the proposed methodology, demonstrating its robustness and accuracy in practical settings.