double machine learning

Design and implement estimation procedures for a low-dimensional target parameter that combine flexible machine learners for high-dimensional nuisance functions with orthogonalization/residualization and cross-fitting to remove first-order bias and enable valid, debiased inference. Build workflows that estimate and bias-correct the target (using orthogonal score constructions and residualisation), and run diagnostics and sensitivity checks across different nuisance learners to ensure robustness to nuisance estimation errors.

doublemachinelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$169K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This paper addresses the problem that estimation errors in high-dimensional nuisance parameters contaminate inference on target causal parameters in semiparametric models. To mitigate this, it proposes and systematically develops the double/debiased machine learning (DML) framework. Built upon Neyman orthogonality and cross-fitting, DML substantially reduces sensitivity to misspecification of nuisance parameter models, thereby enhancing robustness and asymptotic efficiency of target parameter estimation. A key contribution is the principled integration of flexible machine learning methods—including Lasso, random forests, and neural networks—into semiparametric inference, while preserving √n-consistency and asymptotic normality, and enabling modeling of complex data such as text. Empirical results demonstrate substantial improvements in confidence interval coverage and hypothesis testing power. The framework thus provides a unified, interpretable, and reliable methodology for causal inference in high-dimensional and nonlinear settings.

Applying DML to complex data types like text dataIntroducing Neyman orthogonality and cross-fitting in DMLReducing impact of nuisance parameters on target parameter estimation

This work addresses the suboptimal inference in semiparametric estimation caused by estimation errors in nuisance functions when using black-box machine learning models. The authors propose a novel estimator that, without imposing additional assumptions, eliminates first-order stochastic errors from nuisance estimation and achieves optimal convergence rates even when auxiliary functions cannot be consistently estimated. Built upon the framework of orthogonal scores and semiparametric linear functionals, the proposed estimator attains the sharp rate \(n^{-1/2} + \delta^a_\mu + (\delta^s_\mu)^2\) and is shown to be asymptotically normal with minimal asymptotic variance. Its tuning strategy favors undersmoothing and substantially outperforms classical double machine learning methods, making it well-suited for widespread applications such as average treatment effect estimation.

bias-variance trade-offblack-box modelsdouble machine learning

Perturbed Double Machine Learning: Nonstandard Inference Beyond the Parametric Length

Nov 02, 2025
MZ
Mengchu Zheng
🏛️ Rutgers, The State University of New Jersey | Zhejiang University

This paper addresses robust inference for a low-dimensional target parameter β in the presence of infinite-dimensional nuisance parameters. Conventional double machine learning (DML) requires nuisance estimators to converge at a rate faster than n⁻¹/⁴; otherwise, Wald-type confidence intervals become invalid. To overcome this limitation, we propose Perturbed Double Machine Learning (P-DML): it injects random perturbations into nuisance estimation to generate multiple β estimates, then applies bias screening and thresholding to retain only valid ones—thereby relaxing the n⁻¹/⁴ rate requirement. P-DML accommodates both Lasso and general machine learning algorithms without imposing strong convergence conditions on nuisance estimators. We establish that the resulting estimator possesses oracle properties and yields asymptotically valid confidence intervals. Simulation studies demonstrate that P-DML maintains nominal coverage even when nuisance parameters converge slowly, substantially enhancing the robustness of statistical inference.

Inference on low-dimensional parameters with infinite-dimensional nuisancesRobust coverage for DML estimators using perturbation and filteringValid inference when nuisance estimators converge slower than n^{-1/4}

Average partial effect estimation using double machine learning

Aug 17, 2023
HK
Harvey Klyne
🏛️ University of Cambridge

This paper addresses robust estimation of average partial effects (APEs) in nonlinear models under moderate-dimensional settings. We propose a novel double machine learning framework that dispenses with linearity assumptions and differentiability requirements on the regression model, permitting arbitrary black-box machine learning algorithms as first-stage estimators. Our method innovatively introduces re-smoothing to confer differentiability upon otherwise non-differentiable estimators; integrates a location-scale model to flexibly characterize the conditional distribution of covariates; and constructs a doubly robust semiparametric inference procedure. We establish theoretical guarantees: the estimator achieves the semiparametric efficiency bound and remains robust under model misspecification and other nonstandard conditions. Numerical experiments demonstrate substantial improvements over existing APE estimators in both estimation accuracy and confidence interval coverage.

Controls errors via mean and standard deviation estimationEstimates average partial effects interpretably in regressionOvercomes non-differentiability in machine learning methods

The Fundamental Limits of Structure-Agnostic Functional Estimation

May 06, 2023
SB
Sivaraman Balakrishnan
🏛️ Carnegie Mellon University

This paper addresses the problem of estimating functionals of an unknown target function under a structure-agnostic setting—where no specific structural assumptions (e.g., Hölder smoothness) are imposed on the nuisance function, and only a generic convergence rate for nuisance estimation is assumed. Methodologically, it introduces the first formal framework for structure-agnostic estimation, operating under three simultaneous constraints: weak regularity conditions, compatibility with general-purpose nuisance estimators, and sample splitting. Theoretically, it establishes, for the first time, the essential optimality of first-order debiased estimators in this setting. Through minimax lower bound analysis, higher-order perturbation theory, and a unified debiasing framework, the paper precisely characterizes the optimal convergence rate and quantifies the fundamental trade-off between incorporating structural priors and improving estimation efficiency. These results provide foundational theoretical support for nonparametric and semiparametric inference.

Highlights tradeoffs between agnostic and structure-aware estimationInvestigates limits of structure-agnostic functional estimationShows first-order debiasing methods are optimal under weak conditions

Latest Papers

What's happening recently
View more

Sharp Structure-Agnostic Lower Bounds for General Functional Estimation

Dec 19, 2025
JJ
Jikai Jin
🏛️ Stanford University

This paper investigates the minimax optimal error lower bounds for structure-agnostic (i.e., model-agnostic) nonparametric functional estimation, focusing on causal parameters such as the average treatment effect (ATE) and general functional targets. Methodologically, it integrates double machine learning (DML), first-order debiasing, double robustness analysis, and minimax lower bound theory. The key contributions are: (i) the first systematic characterization of structure-agnostic optimal convergence rates under both doubly robust and non-doubly robust settings; (ii) a rigorous proof that DML achieves the theoretical minimax optimal rate across all structure-agnostic scenarios, with explicit closed-form expressions for the optimal rates; and (iii) unification and generalization of existing ATE lower bound results, thereby establishing the universal minimax optimality of DML in this framework. These findings provide foundational theoretical support for model-free causal and statistical inference.

Determines when double robustness is achievable for general nuisance functionalsEstablishes optimal error rates for structure-agnostic functional estimationValidates debiased machine learning as optimal without structural assumptions

This study addresses the challenge of constructing Neyman-orthogonal scores for robust causal inference in semiparametric models with infinite-dimensional nuisance parameters. The authors propose a general framework that, for the first time, explicitly constructs orthogonal scores for a broad class of such models, yielding estimators of the target parameter that are asymptotically normal and require only a convergence rate of $o_p(n^{-1/4})$ or better for the nuisance parameter estimates. The approach seamlessly integrates with machine learning algorithms and is applied to estimate causal effects under binary instrumental variables. Numerical experiments demonstrate substantial finite-sample improvements over naive estimators, and an empirical analysis of the Oregon Health Insurance Experiment confirms the method’s robustness and practical utility in real-world settings.

causal inferenceinstrumental variableNeyman-orthogonal scores

This study addresses the limitations of standard first-order semiparametric estimators in causal inference and missing data problems, which often fail to achieve asymptotic efficiency due to slow convergence of the nuisance functions and exhibit poor finite-sample performance. The authors systematically compare three classes of higher-order efficient estimators—Higher-Order Influence Functions (HOIF), kernel-based HOTMLE, and HAL-HOTMLE—evaluating, for the first time within a unified simulation framework, how their higher-order expansion constructions and regularization strategies affect estimation accuracy. Results demonstrate that higher-order debiasing substantially reduces bias, with HAL-HOTMLE showing robust performance, whereas HOIF proves sensitive to basis truncation and tuning parameters. The work clarifies the conditions under which higher-order corrections are effective in both theory and practice, while highlighting their limitations and key trade-offs for method selection.

approximation strategiesfinite-sample performancehigher-order efficient estimators

This work addresses the challenges of achieving efficient semiparametric estimation of low-dimensional target parameters in the presence of high-dimensional nuisance functions, where out-of-sample prediction of nuisances and statistical efficiency are critical concerns. The authors propose a general, estimator-agnostic cross-fitting engine that automatically executes reproducible folding schedules based on user-specified target functionals and a directed acyclic graph (DAG) encoding the nuisance model structure. Innovatively leveraging graph-based nuisance modeling, the framework supports folding strategies such as disjointness and independence enhancement, while reducing branching dependencies through node replication. It further incorporates explicit scheduling, dependency validation, intelligent caching, and fault isolation mechanisms. Implemented as a lightweight R package publicly available on CRAN, this approach substantially improves the controllability, reproducibility, and computational efficiency of cross-fitting estimators, facilitating large-scale simulations and rapid prototyping.

cross-fittingdouble machine learningfold allocation

This study addresses the regularization bias and overfitting commonly induced by high-dimensional nuisance parameters in target parameter estimation. The authors propose a general debiased machine learning framework that identifies the target parameter through moment conditions and corrects the bias introduced by machine learning estimates of high-dimensional nuisance components via the Riesz representer. They innovatively show that the Riesz representer is uniquely determined by a quadratic functional and incorporate shape constraints into the nonlinear parameter space to mitigate the curse of dimensionality and enhance estimation accuracy. The method combines sieve techniques with deep neural networks to estimate nuisance functions while jointly integrating shape constraints and Riesz representer estimation. Simulation and empirical results demonstrate that the proposed framework achieves unbiased and more precise parameter estimates in high-dimensional settings.

Debiased Machine LearningHigh-dimensional NuisanceMoment Condition

Hot Scholars

MH

Martin Huber

University of Fribourg
microeconometricstreatment evaluationempirical labor economicsempirical health economics
YC

Yifan Cui

Zhejiang University
StatisticsInferenceLearning
FL

Fan Li

Department of Statistical Science, Duke University
statisticscausal inferencecomparative effectiveness researchmissing data
DF

Dennis Frauen

PhD student, LMU Munich
Machine LearningCausal inferenceStatistics
SF

Stefan Feuerriegel

Professor, LMU Munich
AI in ManagementBusiness AnalyticsComputational Social ScienceAI for Good