Score
Design, implement, and analyze regression-adjusted estimators that incorporate efficient influence functions to provide doubly robust bias correction and attain semiparametric efficiency; this entails constructing EIF-based adjustment terms that combine outcome-regression and other nuisance estimates. Analyze the estimators' asymptotic behavior and finite-sample error, including first- and sharp second-order rates under standard nuisance-rate conditions and extensions to batched or grouped data settings.
In causal inference, machine learning is widely used to estimate propensity scores and outcome models, yet practical guidance for constructing target minimum loss estimators (TMLEs) that balance theoretical rigor with implementability remains scarce for applied researchers. Method: Building on the efficient influence function framework, we propose a modular, reproducible TMLE construction procedure: first fit auxiliary models using arbitrary machine learning methods; then perform a one-dimensional targeted update to correct bias. Contribution/Results: Our key innovation lies in translating abstract efficiency theory into three intuitive steps—initial estimation, efficient influence function computation, and parameter update—substantially lowering the conceptual and implementation barriers to TMLE. The estimator retains double robustness and asymptotic efficiency under model misspecification. By bridging the critical gap between statistical theory and empirical practice, our approach enables non-statisticians to conduct reliable, machine learning–enhanced causal analyses.
This paper addresses robust estimation of average partial effects (APEs) in nonlinear models under moderate-dimensional settings. We propose a novel double machine learning framework that dispenses with linearity assumptions and differentiability requirements on the regression model, permitting arbitrary black-box machine learning algorithms as first-stage estimators. Our method innovatively introduces re-smoothing to confer differentiability upon otherwise non-differentiable estimators; integrates a location-scale model to flexibly characterize the conditional distribution of covariates; and constructs a doubly robust semiparametric inference procedure. We establish theoretical guarantees: the estimator achieves the semiparametric efficiency bound and remains robust under model misspecification and other nonstandard conditions. Numerical experiments demonstrate substantial improvements over existing APE estimators in both estimation accuracy and confidence interval coverage.
In causal mediation analysis, conventional estimators of the mediated effect functional suffer from low accuracy and high sensitivity to misspecification of nuisance functions. To address this, we propose a bias-structure-guided two-stage framework that decouples nuisance function estimation. In Stage I, we estimate only the bias-relevant component of the mediation mechanism—rather than the full mechanism—thereby reducing model dependence. In Stage II, we introduce a nonparametric weighted balancing estimator, where weights are constructed by directly optimizing the asymptotic bias of the mediated effect estimator. We establish theoretical guarantees: the resulting estimator is consistent and asymptotically normal, and remains robust under partial misspecification of nuisance functions. Compared with standard approaches, our method substantially improves estimation accuracy and reliability. It provides a principled tool for mediation inference in high-dimensional settings or under model uncertainty.
Accurate estimation and inference for the average treatment effect (ATE) in multi-covariate randomized controlled trials (RCTs) remain challenging, particularly under high-dimensional covariates and small sample sizes. Method: Building on the Neyman finite-population framework, we propose a bias-corrected regression-adjustment estimator with cross-fitting and introduce, for the first time, an HC3-type heteroskedasticity-robust standard error tailored to high-dimensional settings. We rigorously derive first- and second-order stochastic expansions of the random component of regression-adjustment estimators, identifying the source of higher-order bias in conventional inference; leveraging this insight, we design a cross-fitting procedure to eliminate bias and extend HC3 standard errors to stratified experimental designs. Results: Simulations and reanalysis of Angrist et al. (2009)’s education RCT demonstrate that our method substantially improves estimation accuracy, confidence interval coverage, and statistical power—especially in small-sample and high-dimensional scenarios.
To address the low precision of average treatment effect (ATE) estimation in randomized controlled trials (RCTs) with high-dimensional covariates (p ≫ n), this paper proposes a novel covariate adjustment method based on higher-order influence functions (HOIFs). The method systematically establishes the theoretical advantages of HOIFs in RCTs for the first time, unifies a broad class of state-of-the-art adjusted estimators, and rigorously characterizes the conditions under which HOIF-based estimation strictly dominates both unadjusted and linear-model-adjusted estimators. We prove that the proposed estimator achieves semiparametric asymptotic efficiency—i.e., it attains the semiparametric efficiency bound under mild regularity conditions. Numerical simulations and empirical analyses demonstrate substantial gains in estimation accuracy and robustness when p is large relative to n. An accompanying R package, implementing the method, has been publicly released on CRAN.
This study addresses the estimation of the local average treatment effect for the treated (LATT) in an Instrumented Difference-in-Differences (IDiD) setting with covariates and staggered instrument exposure. It derives, for the first time, the efficient influence function (EIF) for LATT under the IDiD framework and constructs a doubly robust estimator applicable to both panel and repeated cross-sectional data, explicitly distinguishing between never-exposed and not-yet-exposed control groups. The proposed method integrates cross-fitting and double machine learning to accommodate high-dimensional covariates while preserving double robustness and strong finite-sample performance. Theoretical analysis shows that under one-sided compliance and absorbing treatment, the LATT can be expressed as a convex combination of the group-time average treatment effects ATT(g,t) introduced by Callaway and Sant’Anna (2021). An implementation of the method is available in the Python package idid.
This work addresses the suboptimal inference in semiparametric estimation caused by estimation errors in nuisance functions when using black-box machine learning models. The authors propose a novel estimator that, without imposing additional assumptions, eliminates first-order stochastic errors from nuisance estimation and achieves optimal convergence rates even when auxiliary functions cannot be consistently estimated. Built upon the framework of orthogonal scores and semiparametric linear functionals, the proposed estimator attains the sharp rate \(n^{-1/2} + \delta^a_\mu + (\delta^s_\mu)^2\) and is shown to be asymptotically normal with minimal asymptotic variance. Its tuning strategy favors undersmoothing and substantially outperforms classical double machine learning methods, making it well-suited for widespread applications such as average treatment effect estimation.
This work addresses the suboptimal convergence rates of classical influence functions when estimating complex, implicitly defined causal parameters such as quantile treatment effects. While existing higher-order methods are limited to explicitly defined parameters, this paper extends the higher-order influence function framework to implicit M- and Z-estimation problems for the first time. By integrating U-process theory with nonparametric estimation, the authors construct a debiased estimator that substantially relaxes the stringent Hölder smoothness assumptions typically imposed on nuisance parameters. The proposed approach achieves improved convergence rates in settings like quantile treatment effect estimation and reduces requirements on model complexity, thereby broadening the applicability of higher-order influence function methodology to a wider class of semiparametric problems.
This study addresses the quantification of omitted variable bias in nonlinear instrumental variable (IV) estimation by extending sensitivity analysis to nonlinear IV frameworks, encompassing local average treatment effects (LATE), LATE for treated individuals (LATT), and partially linear IV models (PLIVM). The authors derive bias decompositions, construct partial identification bounds, and develop computable bias bounds alongside robust inference procedures that adjust confidence intervals accordingly. Integrating double machine learning (DML), the approach accommodates flexible control for high-dimensional covariates. Application to the JTPA experiment reveals that estimated program effects for women remain robustly significant, whereas those for men are sensitive to potential omitted variables; first-stage compliance rate estimates are stable, but intent-to-treat and treatment effect estimates exhibit greater fragility.
This study addresses the limitations of standard first-order semiparametric estimators in causal inference and missing data problems, which often fail to achieve asymptotic efficiency due to slow convergence of the nuisance functions and exhibit poor finite-sample performance. The authors systematically compare three classes of higher-order efficient estimators—Higher-Order Influence Functions (HOIF), kernel-based HOTMLE, and HAL-HOTMLE—evaluating, for the first time within a unified simulation framework, how their higher-order expansion constructions and regularization strategies affect estimation accuracy. Results demonstrate that higher-order debiasing substantially reduces bias, with HAL-HOTMLE showing robust performance, whereas HOIF proves sensitive to basis truncation and tuning parameters. The work clarifies the conditions under which higher-order corrections are effective in both theory and practice, while highlighting their limitations and key trade-offs for method selection.