instrumental variables

Design and implement estimation and testing procedures that use instrumental variables to identify causal effects in the presence of endogeneity, including two-stage methods (2SLS, 2SRI/2SRI), IV estimation and inference, selection and testing of instrument relevance and exogeneity, and recovery of local average treatment effects. Extend and apply these estimators and inferential corrections to specialized outcomes such as time-to-event/survival data using martingale-residual–based estimating equations, account for propagated uncertainty, and perform diagnostics and sensitivity analyses for instrument validity.

instrumentalvariables

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses causal effect identification in observational settings with unmeasured confounding and potentially invalid instrumental variables, focusing on linear instrumental variable models with multiple endogenous treatments. The authors propose generalized majority and plurality rules to achieve identification, coupled with a data-driven instrument selection procedure that yields sampling confidence intervals robust to the erroneous inclusion of invalid instruments. Under standard regularity conditions, these intervals are shown to attain asymptotic nominal coverage and exhibit length shrinking at the parametric rate. The practical utility and validity of the proposed method are demonstrated through an empirical application in Mendelian randomization.

causal inferenceinstrumental variablesinvalid instruments

Inference for Heterogeneous Treatment Effects with Efficient Instruments and Machine Learning

Mar 05, 2025
CS
Cyrill Scheidegger
🏛️ ETH Zurich | Rutgers University

Estimating heterogeneous treatment effects under endogeneity remains challenging, particularly when instrumental variables (IVs) are weak. Method: This paper proposes a novel semiparametric IV estimation framework that integrates double/debiased machine learning (DML), machine learning–based IV estimation (MLIV), and kernel smoothing. It is the first to embed MLIV within the DML architecture to construct confidence sets robust to weak instruments. Contribution/Results: We establish consistency and asymptotic normality of the estimator and provide unified statistical inference guarantees. Implemented in the R package `IVDML`, the method demonstrates substantial improvements in confidence interval coverage and estimation accuracy under weak-IV settings, both on synthetic and real-world data. Our approach offers a new paradigm for causal heterogeneity analysis—rigorous in theory and feasible in computation.

Developing robust confidence sets for weak instrumental variable scenariosEstimating heterogeneous treatment effects with endogeneity using instrumental variablesProviding accessible implementation in R package for practical application

Learning control variables and instruments for causal analysis in observational data

Jul 05, 2024
NA
Nicolas Apfel
🏛️ University of Innsbruck | University of York | University of Fribourg | Heinrich Heine University Düsseldorf

Estimating causal effects from observational data requires selecting appropriate control and instrumental variables that satisfy causal identification conditions—a challenging task often reliant on strong domain knowledge or ad hoc assumptions. Method: This paper proposes the first end-to-end joint learning framework that automatically identifies valid combinations of control and instrumental variables. Grounded in conditional independence testing, the method integrates nonparametric dependence measures with structural search optimization, ensuring statistical consistency in variable selection under mild regularity conditions. Contribution/Results: Unlike conventional approaches requiring prespecified variable sets or strong prior assumptions, our framework is fully data-driven. In simulations, it achieves significantly higher variable identification accuracy. Empirically, applied to the Job Corps study, its estimated treatment effect closely aligns with results from the randomized controlled trial—demonstrating both validity and robustness in real-world causal inference.

Detects control variables and instruments for causal analysis in observational dataLearns partition of instruments and control variables from observed dataTests joint existence of instruments and control variables using machine learning

Causal Effect Identification and Inference with Endogenous Exposures and a Light-tailed Error

Aug 12, 2024
RW
Ruoyu Wang
🏛️ Harvard T.H. Chan School of Public Health | Peking University

Endogeneity in exposures impedes causal identification, and conventional approaches rely either on strong functional-form assumptions or valid instrumental variables (IVs). This paper proposes an extremal conditional quantile contrast method grounded in a light-tailed error assumption. We establish, for the first time, that extreme quantile regression is inherently robust to endogeneity under light-tailed errors—enabling causal identification without IVs or additional parametric restrictions. Theoretically, we prove strong consistency of the estimator and asymptotic normality in linear models. Simulation studies and empirical analysis using automobile sales data demonstrate that the method maintains high estimation accuracy and reliable confidence interval coverage even when invalid IVs are present. By circumventing reliance on external instruments or stringent modeling assumptions, our approach substantially broadens the scope of applicable settings for endogeneity-robust inference.

Address endogeneity in causal inference with invalid auxiliary variablesEstimate causal effects using extreme quantile regression without auxiliary variablesIdentify causal effects with endogenous exposures and light-tailed errors

Inference for Treatment Effects Conditional on Generalized Principal Strata using Instrumental Variables

Nov 07, 2024
YB
Yuehao Bai
🏛️ University of Southern California | University of Chicago | Massachusetts Institute of Technology | University of California–Los Angeles | Yale University

This paper addresses conditional treatment effect inference based on generalized principal strata—defined as response-type vectors—under multivalued treatments, multivalued outcomes, and instrumental variables (IVs). First, it systematically characterizes the identification region for such effects, equivalently reformulating it as the existence problem of solutions to a linear system subject to nonnegativity and structural constraints. Leveraging IV exogeneity and the zero-probability assumption on certain response types, the paper proposes a unified inferential framework grounded in linear programming and convex optimization—extending Fang et al. (2023). It further develops a computationally tractable, conservative, and consistent procedure for constructing confidence sets applicable to canonical causal parameters, including the population stratification effect (PSE) and variants of the complier average causal effect (CACE). The approach substantially improves inference precision and broadens applicability across complex multivalued settings.

Characterizing identified sets under exogeneity and response restrictionsEstimating treatment effects conditional on generalized principal strataUsing instrumental variables for discrete treatments and outcomes

Latest Papers

What's happening recently
View more

This study addresses the limitation of conventional instrumental variable (IV) methods, which often assume a constant treatment effect and thus struggle to accommodate effect heterogeneity in real-world settings. Building on the local average treatment effect (LATE) framework, the paper systematically integrates covariate-adjusted IV approaches, clarifying how covariates influence the weighting structure of LATE estimators. It proposes flexible modeling strategies to avoid parametric misspecification and incorporates robust diagnostic tests for violations of the monotonicity assumption. By combining nonparametric and semiparametric estimation techniques, formal hypothesis testing, and accompanying software implementation, this work offers empirical researchers a theoretically rigorous yet practically feasible causal inference workflow, substantially enhancing the reliability and applicability of IV analysis.

Causal InferenceEmpirical PracticeHeterogeneous Treatment Effects

This study addresses the estimation bias arising from endogeneity in regression models by proposing a general and computationally efficient semiparametric projection method. The approach constructs endogenous instrumental variables by projecting and expanding the conditional mean function of the structural error onto the space of explanatory variables, thereby avoiding reliance on conventional exogenous instruments or specific parametric model forms. It is applicable to linear, nonlinear, and semiparametric settings alike. By integrating LASSO-based variable selection with asymptotic theory, the paper establishes identification conditions and asymptotic properties of the resulting estimator. Extensive simulations and empirical analyses demonstrate the method’s strong finite-sample performance, confirming its practical utility in mitigating endogeneity bias across diverse modeling contexts.

endogeneityidentificationinstrumental variables

This study addresses the challenge of identifying causal effects in static panel data when the treatment is endogenous and high-dimensional nonlinear confounders are present, a setting where conventional instrumental variable (IV) methods often fail. The paper proposes Panel IV-DML, the first extension of double machine learning (DML) to a panel IV framework, which integrates flexible machine learning techniques—such as Lasso and random forests—for covariate adjustment and introduces a novel weak identification diagnostic tailored to this setting. Theoretical analysis and Monte Carlo simulations demonstrate that the estimator achieves higher precision under strong instruments and more robust inference under weak instruments. Empirical applications across three immigration studies confirm that the method replicates classic 2SLS findings while also detecting scenarios of weak identification, thereby supporting more cautious causal conclusions.

causal inferenceendogeneityhigh-dimensional confounding

A Synthetic Instrumental Variable Method: Using the Dual Tendency Condition for Coplanar Instruments

Dec 19, 2025
RD
Ratbek Dzhumashev
🏛️ Monash University | Data61, CSIRO

Traditional instrumental variable (IV) methods often suffer from causal bias due to weak or invalid instruments and reliance on external data. To address this, we propose a novel data-driven approach that constructs synthetic instrumental variables (SIVs) solely from observed covariates. Our method introduces the “double-tilting (DT) condition”—a newly established identification criterion that enables valid IV selection without external instruments and further determines the sign of the correlation between the endogenous variable and the structural error. By integrating DT-condition testing with heteroskedasticity-robust estimation, our framework substantially improves causal effect estimation accuracy in both simulations and empirical applications. It effectively mitigates weak instrument and instrument invalidity issues while drastically reducing dependence on exogenous instruments. This work establishes a verifiable, purely observational paradigm for addressing endogeneity, advancing causal inference methodology beyond conventional IV assumptions.

Constructs valid instruments using only existing dataIdentifies valid instruments without requiring external variablesImproves causal inference by mitigating common IV limitations

Hot Scholars

VS

Vasilis Syrgkanis

Assistant Professor, Stanford University
Machine LearningCausal InferenceEconometricsGame Theory
ZL

Zhonghua Liu

Department of Biostatistics, Columbia University
StatisticsMachine LearningCausal InferenceGenetics/Genomics
YC

Yifan Cui

Zhejiang University
StatisticsInferenceLearning
VS

Vasilis Sarafidis

Brunel University London
econometricspanel data analysisspatial econometrics
UB

Ufuk Beyaztas

Department of Statistics, Marmara University
Functional data analysisResampling methodsTime series analysisStatistics