jackknife resampling

Applying jackknife-based resampling estimators to reduce bias and quantify uncertainty (e.g., for Sobol' index estimation), deriving their mean-squared-error properties under Monte Carlo and adjusting inference for selection effects.

jackkniferesampling

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses bias and efficiency issues in estimating Sobol’ indices for global sensitivity analysis by unifying classical pick-freeze and nested Monte Carlo estimators within a nested simulation framework, revealing that the former is a special case with fixed inner-sample size. Building on this insight, the authors propose two novel jackknife-based estimators—unbiased jackknife and split-sample jackknife—and provide the first systematic evaluation of how Latin hypercube sampling (LHS) affects bias correction across different estimators. Theoretical analysis demonstrates that the split-sample jackknife estimator achieves the standard mean squared error (MSE) convergence rate under a fixed computational budget, outperforming conventional approaches. Numerical experiments corroborate the theoretical findings and offer practical guidance for selecting estimators in real-world applications.

bias correctioncomputational budgetnested simulation

This study addresses the lack of an automatic, efficient, and model-agnostic inference procedure for fixed-effects models. The authors propose a general inference framework that constructs a self-normalized jackknife t-statistic from subsample estimators, enabling direct computation of hypothesis tests, confidence intervals, and p-values. The method requires no complex tuning parameters, offers high computational efficiency, and exhibits strong robustness to model specification. Applicable across a broad class of fixed-effects models, this approach substantially streamlines statistical inference and facilitates automated, unified inferential practice.

fixed effectsinferencejackknife

Traditional bootstrap methods are computationally expensive and lack theoretical guarantees in semi-parametric and machine learning settings, undermining inference reliability. This work proposes V-fold jackknife as an efficient alternative: by refitting the estimator only V times on leave-one-fold-out samples, it quantifies uncertainty directly through the empirical variance of jackknife pseudovalues, bypassing explicit influence function calculations. For the first time, the authors establish a t-distribution limit theory for Studentized V-fold jackknife statistics with fixed V and extend it to generalized asymptotically linear estimators, enabling scale-invariant inference without requiring knowledge of convergence rates. The method demonstrates robust performance across diverse applications—including average treatment effects, Kaplan–Meier survival curves, and highly adaptive Lasso dose–response curves—and remains reliable even when influence functions fail to exist or are difficult to estimate.

bootstrapconfidence intervalsjackknife

This paper addresses inference for parameters defined by multivariate three-sample U-statistics—such as the volume under the ROC surface (VUS) in multi-class classification—where conventional asymptotic approximations suffer from poor coverage and high computational cost. We propose a jackknife pseudo-value-based empirical likelihood (JEL) framework, the first to extend empirical likelihood to multivariate three-sample U-statistics. Under mild regularity conditions, the proposed JEL ratio statistic follows a Wilks-type chi-square asymptotic distribution, obviating explicit variance estimation and computationally intensive resampling. Monte Carlo simulations demonstrate that our method achieves coverage probabilities closer to nominal levels than normal approximation or kernel-based approaches, while reducing computation time substantially. Empirical evaluation on real classification datasets confirms its robustness and practical utility. Key contributions include: (i) theoretical innovation—a novel inferential framework with rigorous asymptotic theory; (ii) methodological advantages—variance-free and resampling-free inference; and (iii) applied impact—simultaneous improvements in accuracy and efficiency for VUS inference.

Applies method to classification problems using volume under surface measuresConstructs confidence intervals without explicit variance estimation or resamplingDevelops jackknife empirical likelihood for multivariate three-sample U-statistics

Scalable Efficient Inference in Complex Surveys through Targeted Resampling of Weights

Apr 15, 2025
SD
Snigdha Das
🏛️ Texas A&M University | Virginia Commonwealth University | University of Wisconsin-Madison

Traditional inference methods suffer from bias in complex surveys (e.g., stratified, multistage sampling) due to informative sampling, and existing pseudo-likelihood approaches lack finite-sample uncertainty quantification. Bayesian pseudo-posterior and weighted bootstrap methods, respectively, suffer from miscalibrated confidence coverage and computational inefficiency. This paper proposes the Survey-adjusted Weighted Likelihood Bootstrap (S-WLB)—the first survey-weighted likelihood bootstrap method—where the target distribution is constructed using true sampling weights and resampling is performed with corresponding weights. S-WLB ensures asymptotic validity while delivering accurate finite-sample inference. Theoretical analysis establishes rigorous confidence interval guarantees under standard survey design assumptions. Empirical evaluations on NHANES and NSDUH data demonstrate that S-WLB achieves significantly higher computational efficiency than weighted Bayesian bootstrap and yields empirical coverage rates markedly closer to nominal levels.

Addresses biased estimators in complex survey designsImproves finite-sample uncertainty quantification methodsOvercomes computational limitations of Bayesian approaches

Latest Papers

What's happening recently
View more

This study addresses the challenge of accurately determining effective degrees of freedom for variance estimation in stratified sampling designs with only two primary sampling units per stratum—a limitation that undermines the reliability of confidence intervals. By examining the between-stratum independence of variance components in both Balanced Repeated Replication (BRR) and paired jackknife methods, the authors develop a unified analytical framework. Leveraging the orthogonality of Hadamard matrices, they decompose BRR into independent between-stratum contrast terms and, in conjunction with the Welch–Satterthwaite approximation, derive explicit formulas for the effective degrees of freedom under both approaches. This advancement substantially enhances the accuracy and applicability of confidence intervals for population totals under complex sampling schemes.

balanced repeated replicationeffective degrees of freedompaired jackknife

In finite-population causal inference, interference effects preclude the existence of a universally applicable and conservative variance estimator. This work proposes the Neyman Jackknife framework, which achieves conservative inference for the variance of causal effect estimators under arbitrary interference structures by re-estimating treatment effects after omitting subsets of treatment assignments. The approach unifies the classical Neyman variance estimator under the Stable Unit Treatment Value Assumption (SUTVA) with the Newey-West heteroskedasticity-and-autocorrelation-consistent (HAC) estimator from time series analysis, offering a general and flexible blueprint for variance estimation. Numerical experiments demonstrate that the proposed framework performs comparably to, and often better than, specialized baseline methods across a range of interference settings.

causal inferencefinite populationinterference

This study addresses linear regression models plagued by numerous potentially weak instrumental variables, endogeneity, and heteroskedasticity. It proposes a novel class of jackknife-based test statistics for hypotheses concerning the full parameter vector or general linear restrictions. Constructed within the generalized method of moments framework, the proposed tests exhibit a chi-square mixture as their asymptotic null distribution; by appropriately adjusting the objective function, they attain a standard chi-square limit, substantially enhancing inferential practicality and interpretability. Both theoretical analysis and Monte Carlo simulations demonstrate superior finite-sample performance, notably outperforming Anderson–Rubin–type procedures. The methodology is successfully applied to UK Biobank data, uncovering a causal effect of alcohol consumption on body mass index (BMI).

endogeneityheteroskedasticityhypothesis testing

This study addresses the inefficiency and instability of traditional global sensitivity analysis methods under small sample sizes. The authors propose a unified nonparametric estimation framework based on rank statistics and Chatterjee’s recently introduced empirical correlation coefficient, extending it for the first time to diverse sensitivity indices—including Cramér–von Mises, first-order Sobol’, metric space–based measures, and higher-order moments. Theoretical analysis establishes the consistency and asymptotic normality of the proposed estimators. Numerical experiments demonstrate that the method achieves superior computational efficiency and stability in small-sample settings, offering a theoretically grounded and practically effective new tool for global sensitivity analysis.

Cramér-von-Mises indicesGlobal Sensitivity Analysisrank statistics

Hot Scholars

AE

Ashkan Ertefaie

Associate Professor, University of Rochester
Causal inferenceDynamic treatment regimesSemiparametric theorySurvival analysis
MV

Mark van der Laan

Jiann-Ping Hsu/Karl E. Peace Professor of Biostatistics & Statistics, University of California Berkeley
StatisticsBiostatisticsCausal InferenceMachine Learning