Double Machine Learning for Static Panel Data with Instrumental Variables: New Method and Applications

📅 2026-03-20
📈 Citations: 0
Influential: 0
📄 PDF

career value

202K/year
🤖 AI Summary
This study addresses the challenge of identifying causal effects in static panel data when the treatment is endogenous and high-dimensional nonlinear confounders are present, a setting where conventional instrumental variable (IV) methods often fail. The paper proposes Panel IV-DML, the first extension of double machine learning (DML) to a panel IV framework, which integrates flexible machine learning techniques—such as Lasso and random forests—for covariate adjustment and introduces a novel weak identification diagnostic tailored to this setting. Theoretical analysis and Monte Carlo simulations demonstrate that the estimator achieves higher precision under strong instruments and more robust inference under weak instruments. Empirical applications across three immigration studies confirm that the method replicates classic 2SLS findings while also detecting scenarios of weak identification, thereby supporting more cautious causal conclusions.

Technology Category

Application Category

📝 Abstract
Panel data methods are widely used in empirical analysis to address unobserved heterogeneity, but causal inference remains challenging when treatments are endogenous and confounding variables high-dimensional and potentially nonlinear. Standard instrumental variables (IV) estimators, such as two-stage least squares (2SLS), become unreliable when instrument validity requires flexibly conditioning on many covariates with potentially non-linear effects. This paper develops a Double Machine Learning estimator for static panel models with endogenous treatments (panel IV DML), and introduces weak-identification diagnostics for it. We revisit three influential migration studies that use shift-share instruments. In these settings, instrument validity depends on a rich covariate adjustment. In one application, panel IV DML strengthens the predictive power of the instrument and broadly confirms 2SLS results. In the other cases, flexible adjustment makes the instruments weak, leading to substantially more cautious causal inference than conventional 2SLS. Monte Carlo evidence supports these findings, showing that panel IV DML improves estimation accuracy under strong instruments and delivers more reliable inference under weak identification.
Problem

Research questions and friction points this paper is trying to address.

panel data
endogeneity
instrumental variables
high-dimensional confounding
causal inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Double Machine Learning
Panel Data
Instrumental Variables
Weak Identification
High-Dimensional Covariates
A
Anna Baiardi
Erasmus School of Economics, Erasmus University and Tinbergen Institute, Rotterdam, Netherlands
P
Paul S. Clarke
Erasmus School of Economics, Erasmus University and Tinbergen Institute, Rotterdam, Netherlands
A
Andrea A. Naghi
Department of Business Analytics and Applied Economics, Queen Mary University of London, UK
A
Annalivia Polselli
Institute for Social and Economic Research, University of Essex, Colchester, UK