Score
Design and implement jackknife resampling procedures that compute leave-one-out estimators and produce estimates of bias and variance; apply jackknife-based bias correction to adjust estimators and assess mean squared error to evaluate estimator performance.
Traditional bootstrap methods are computationally expensive and lack theoretical guarantees in semi-parametric and machine learning settings, undermining inference reliability. This work proposes V-fold jackknife as an efficient alternative: by refitting the estimator only V times on leave-one-fold-out samples, it quantifies uncertainty directly through the empirical variance of jackknife pseudovalues, bypassing explicit influence function calculations. For the first time, the authors establish a t-distribution limit theory for Studentized V-fold jackknife statistics with fixed V and extend it to generalized asymptotically linear estimators, enabling scale-invariant inference without requiring knowledge of convergence rates. The method demonstrates robust performance across diverse applications—including average treatment effects, Kaplan–Meier survival curves, and highly adaptive Lasso dose–response curves—and remains reliable even when influence functions fail to exist or are difficult to estimate.
This study addresses the lack of an automatic, efficient, and model-agnostic inference procedure for fixed-effects models. The authors propose a general inference framework that constructs a self-normalized jackknife t-statistic from subsample estimators, enabling direct computation of hypothesis tests, confidence intervals, and p-values. The method requires no complex tuning parameters, offers high computational efficiency, and exhibits strong robustness to model specification. Applicable across a broad class of fixed-effects models, this approach substantially streamlines statistical inference and facilitates automated, unified inferential practice.
This study addresses bias and efficiency issues in estimating Sobol’ indices for global sensitivity analysis by unifying classical pick-freeze and nested Monte Carlo estimators within a nested simulation framework, revealing that the former is a special case with fixed inner-sample size. Building on this insight, the authors propose two novel jackknife-based estimators—unbiased jackknife and split-sample jackknife—and provide the first systematic evaluation of how Latin hypercube sampling (LHS) affects bias correction across different estimators. Theoretical analysis demonstrates that the split-sample jackknife estimator achieves the standard mean squared error (MSE) convergence rate under a fixed computational budget, outperforming conventional approaches. Numerical experiments corroborate the theoretical findings and offer practical guidance for selecting estimators in real-world applications.
This study addresses the challenge of accurately determining effective degrees of freedom for variance estimation in stratified sampling designs with only two primary sampling units per stratum—a limitation that undermines the reliability of confidence intervals. By examining the between-stratum independence of variance components in both Balanced Repeated Replication (BRR) and paired jackknife methods, the authors develop a unified analytical framework. Leveraging the orthogonality of Hadamard matrices, they decompose BRR into independent between-stratum contrast terms and, in conjunction with the Welch–Satterthwaite approximation, derive explicit formulas for the effective degrees of freedom under both approaches. This advancement substantially enhances the accuracy and applicability of confidence intervals for population totals under complex sampling schemes.
This paper addresses inference for parameters defined by multivariate three-sample U-statistics—such as the volume under the ROC surface (VUS) in multi-class classification—where conventional asymptotic approximations suffer from poor coverage and high computational cost. We propose a jackknife pseudo-value-based empirical likelihood (JEL) framework, the first to extend empirical likelihood to multivariate three-sample U-statistics. Under mild regularity conditions, the proposed JEL ratio statistic follows a Wilks-type chi-square asymptotic distribution, obviating explicit variance estimation and computationally intensive resampling. Monte Carlo simulations demonstrate that our method achieves coverage probabilities closer to nominal levels than normal approximation or kernel-based approaches, while reducing computation time substantially. Empirical evaluation on real classification datasets confirms its robustness and practical utility. Key contributions include: (i) theoretical innovation—a novel inferential framework with rigorous asymptotic theory; (ii) methodological advantages—variance-free and resampling-free inference; and (iii) applied impact—simultaneous improvements in accuracy and efficiency for VUS inference.
This study addresses the issue of over-rejection in traditional difference-in-differences (DiD) and Callaway–Sant’Anna DiD (CSDID) estimators under settings with few clusters, sparse treated clusters, or serial correlation, which can severely distort inference. To mitigate this problem, the paper introduces the cluster jackknife method into the CSDID framework—the first such application—to correct bias arising from cluster structure and substantially improve inference reliability in small samples. By combining CSDID point estimates with jackknife-based standard errors, the proposed approach markedly enhances the accuracy of hypothesis testing in scenarios with a limited number of clusters. Extensive simulations demonstrate its superior finite-sample performance, and the authors provide publicly available software implementations in both Stata (csdidjack) and R (didjack) to facilitate adoption.
This study addresses attenuation bias in scalar-on-density regression when the number of repeated measurements per observational unit is limited, leading to insufficient effective sample size for accurate coefficient function estimation. The work systematically establishes, for the first time, a monotonic decreasing relationship between the number of measurements and the magnitude of attenuation bias. To correct this bias, the authors innovatively integrate the Simulation-Extrapolation (SIMEX) method with bootstrap resampling within a functional data analysis framework, simulating scenarios with fewer measurements and extrapolating estimates to the theoretical limit of infinite replicates. Combining techniques from functional data analysis and density estimation, the proposed approach substantially reduces estimation bias in simulations. Applied to NHANES data, it successfully identifies and corrects finite-measurement bias in the association between physical activity density and all-cause mortality.
In finite-population causal inference, interference effects preclude the existence of a universally applicable and conservative variance estimator. This work proposes the Neyman Jackknife framework, which achieves conservative inference for the variance of causal effect estimators under arbitrary interference structures by re-estimating treatment effects after omitting subsets of treatment assignments. The approach unifies the classical Neyman variance estimator under the Stable Unit Treatment Value Assumption (SUTVA) with the Newey-West heteroskedasticity-and-autocorrelation-consistent (HAC) estimator from time series analysis, offering a general and flexible blueprint for variance estimation. Numerical experiments demonstrate that the proposed framework performs comparably to, and often better than, specialized baseline methods across a range of interference settings.