bootstrap prediction-powered inference

Designs and analyzes estimators and inference procedures that use an auxiliary predictive model f(x) as a control variate to improve efficiency, exploit unlabeled covariates to reduce estimation variance, and produce valid confidence intervals and hypothesis tests while accounting for the predictor’s imperfection. These procedures employ bootstrap resampling or related finite-sample calibration techniques (model-assisted or prediction-powered inference) to obtain reliable uncertainty quantification without relying solely on large-sample asymptotics.

bootstrapprediction-poweredinference

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Quantifying Uncertainty: All We Need is the Bootstrap?

Mar 29, 2024
UZ
Urvsa Zrimvsek
🏛️ University of Ljubljana

This paper addresses the complexity and high pedagogical/practical barriers associated with conventional uncertainty quantification methods—such as standard errors, confidence intervals, and hypothesis tests—in statistical inference. To evaluate the potential of nonparametric bootstrap as a unified alternative, we conduct a large-scale simulation study rigorously comparing single bootstrap, double bootstrap, and classical methods across multiple dimensions: sample size, confidence level, data-generating mechanisms, and statistical functionals. Results demonstrate that the double bootstrap consistently achieves superior coverage accuracy, stability, and robustness—particularly under small-sample and non-normal conditions—outperforming both classical approaches and the single bootstrap. We thus establish the double bootstrap as a principled, parsimonious, and high-performance paradigm for uncertainty quantification, providing both theoretical justification and empirical evidence to support its adoption in statistical education and applied practice.

Assessing bootstrap's potential to simplify statistical education and practiceComparing double bootstrap performance against traditional confidence interval techniquesEvaluating bootstrap as universal alternative for uncertainty quantification methods

Integer Programming for Generalized Causal Bootstrap Designs

Oct 28, 2024
JB
Jennifer Brennan
🏛️ Google Research | University of Southern California | Uber Technologies

In experimental causal inference, design uncertainty—arising from the assignment mechanism—is often overlooked under small sample sizes and heterogeneous treatment effects; conventional causal bootstrap methods apply only to completely randomized designs and average treatment effect estimation. This paper addresses this limitation by introducing integer linear programming into the causal bootstrap framework for the first time, enabling computation of the worst-case copula under generalized assignment mechanisms (e.g., conditional unconfoundedness, bounded confounding) to uniformly calibrate design uncertainty. The method accommodates both linear and quadratic treatment effect estimators and is supported by asymptotic theory establishing its validity. Monte Carlo simulations demonstrate that, in small-scale geographic experiments, the proposed approach substantially improves confidence interval coverage and precision while delivering more robust control of Type I error.

Addresses design uncertainty in small fixed heterogeneous samplesExtends causal bootstrap to non-standard designs and estimatorsGeneralizes method for various assignment types and estimators

This study addresses the challenge of conducting valid Bayesian inference for target parameters—such as causal effects—in the presence of finite-dimensional nuisance parameters. The authors propose a general framework that integrates Bayesian bootstrap with Dirichlet process priors within an estimating equation approach, explicitly accounting for uncertainty in propensity score estimation. The method is robust to model misspecification and remains valid under only single robustness conditions. It extends the "linked Bayesian bootstrap" to nonstandard Bayesian settings, yielding posterior inferences with favorable frequentist properties. Theoretical analysis demonstrates that the resulting posterior distribution exhibits desirable asymptotic behavior, with credible intervals achieving nominal coverage probabilities in large samples.

Bayesian bootstrapcausal inferencenuisance parameter

Sharp variance estimator and causal bootstrap in stratified randomized experiments

Jan 30, 2024
HY
Haoyang Yu
🏛️ Tsinghua University | North Carolina State University | Duke University

In hierarchical randomized experiments, small sample sizes or skewed outcomes often lead to overly conservative Neyman variance estimates and failure of normal approximations, thereby undermining the reliability of inference for weighted average treatment effects. To address this, we propose a randomization-based exact variance estimator and introduce two novel causal bootstrap methods—rank-preserving and constant-treatment-effect counterfactual imputation—achieving second-order accuracy improvements; these are further extended to paired-experiment designs. Our approach is grounded in finite-population asymptotic theory and the randomization-inference framework, and is implemented in the open-source R package *CausalBootstrap*. Simulation studies and empirical applications demonstrate that the proposed methods substantially improve confidence interval coverage and statistical power, effectively mitigating inference bias under small-sample and outcome-skewness conditions.

Address conservativeness and normal approximation failure in Neyman-type estimatorsDevelop causal bootstrap methods for accurate sampling distributionImprove variance estimation in small or skewed stratified experiments

A Note on the Prediction-Powered Bootstrap

May 28, 2024
TZ
Tijana Zrnic
🏛️ Stanford University

Traditional asymptotic theory (e.g., the Central Limit Theorem) often fails for complex research problems, and prediction-driven inference lacks flexible, assumption-light tools. Method: We propose PPBoot—a novel method that directly embeds standard bootstrap resampling into the Prediction-Powered Inference (PPI) framework, eliminating reliance on asymptotic normality assumptions. PPBoot integrates arbitrary black-box prediction models with bootstrap-based uncertainty quantification and incorporates bias correction to enable valid statistical inference for arbitrary estimators. Contribution/Results: Compared to CLT-dependent PPI(++) methods, PPBoot achieves comparable or superior performance in multivariate regression and causal effect estimation. Crucially, it remains robust in settings where the CLT breaks down—such as small samples and non-i.i.d. data—while maintaining conceptual simplicity and ease of implementation. PPBoot thus substantially broadens the applicability and reliability of prediction-powered inference.

Complex ProblemsMathematical TheoryPredictive Research

Latest Papers

What's happening recently
View more

This study addresses Bayesian inference for low-dimensional target parameters in semiparametric models, particularly under the presence of complex nuisance components that may compromise frequentist properties. To this end, we construct posterior distributions by integrating estimating function methods with nonparametric Bayesian techniques—such as Dirichlet processes and Bayesian bootstrap—under conditions weaker than the classical stochastic equicontinuity assumption. We establish asymptotic normality and consistency of the resulting posterior, rigorously identifying the key assumptions required to guarantee desirable frequentist behavior. The theoretical analysis systematically elucidates how relaxing these assumptions affects inferential performance. Extensive simulations corroborate the effectiveness of the proposed methodology, demonstrating its robustness and accuracy in practical settings.

asymptotic normalityBayesian semi-parametric modelsfrequentist properties

Hot Scholars

OS

Osvaldo Simeone

King's College London
Information theorymachine learningquantum information processingwireless systems
CS

Carson Sobolewski

Massachusetts Institute of Technology
Autonomous SystemsControlMachine LearningUncertainty Quantification
AM

Anirban Majumder

Applied Scientist, Amazon Science
Machine LearningDeep LearningNatural Language ProcessingGenerative AI
YC

Yifan Cui

Zhejiang University
StatisticsInferenceLearning
JB

Jonathan Berant

Professor, Tel-Aviv University, Visiting Faculty Researcher, Google DeepMInd
Natural Language ProcessingMachine Learning