Score
Design and implement estimation and testing procedures that use instrumental variables to identify causal effects in the presence of endogeneity, including two-stage methods (2SLS, 2SRI/2SRI), IV estimation and inference, selection and testing of instrument relevance and exogeneity, and recovery of local average treatment effects. Extend and apply these estimators and inferential corrections to specialized outcomes such as time-to-event/survival data using martingale-residual–based estimating equations, account for propagated uncertainty, and perform diagnostics and sensitivity analyses for instrument validity.
This study addresses causal effect identification in observational settings with unmeasured confounding and potentially invalid instrumental variables, focusing on linear instrumental variable models with multiple endogenous treatments. The authors propose generalized majority and plurality rules to achieve identification, coupled with a data-driven instrument selection procedure that yields sampling confidence intervals robust to the erroneous inclusion of invalid instruments. Under standard regularity conditions, these intervals are shown to attain asymptotic nominal coverage and exhibit length shrinking at the parametric rate. The practical utility and validity of the proposed method are demonstrated through an empirical application in Mendelian randomization.
Estimating heterogeneous treatment effects under endogeneity remains challenging, particularly when instrumental variables (IVs) are weak. Method: This paper proposes a novel semiparametric IV estimation framework that integrates double/debiased machine learning (DML), machine learning–based IV estimation (MLIV), and kernel smoothing. It is the first to embed MLIV within the DML architecture to construct confidence sets robust to weak instruments. Contribution/Results: We establish consistency and asymptotic normality of the estimator and provide unified statistical inference guarantees. Implemented in the R package `IVDML`, the method demonstrates substantial improvements in confidence interval coverage and estimation accuracy under weak-IV settings, both on synthetic and real-world data. Our approach offers a new paradigm for causal heterogeneity analysis—rigorous in theory and feasible in computation.
Estimating causal effects from observational data requires selecting appropriate control and instrumental variables that satisfy causal identification conditions—a challenging task often reliant on strong domain knowledge or ad hoc assumptions. Method: This paper proposes the first end-to-end joint learning framework that automatically identifies valid combinations of control and instrumental variables. Grounded in conditional independence testing, the method integrates nonparametric dependence measures with structural search optimization, ensuring statistical consistency in variable selection under mild regularity conditions. Contribution/Results: Unlike conventional approaches requiring prespecified variable sets or strong prior assumptions, our framework is fully data-driven. In simulations, it achieves significantly higher variable identification accuracy. Empirically, applied to the Job Corps study, its estimated treatment effect closely aligns with results from the randomized controlled trial—demonstrating both validity and robustness in real-world causal inference.
Endogeneity in exposures impedes causal identification, and conventional approaches rely either on strong functional-form assumptions or valid instrumental variables (IVs). This paper proposes an extremal conditional quantile contrast method grounded in a light-tailed error assumption. We establish, for the first time, that extreme quantile regression is inherently robust to endogeneity under light-tailed errors—enabling causal identification without IVs or additional parametric restrictions. Theoretically, we prove strong consistency of the estimator and asymptotic normality in linear models. Simulation studies and empirical analysis using automobile sales data demonstrate that the method maintains high estimation accuracy and reliable confidence interval coverage even when invalid IVs are present. By circumventing reliance on external instruments or stringent modeling assumptions, our approach substantially broadens the scope of applicable settings for endogeneity-robust inference.
This paper addresses conditional treatment effect inference based on generalized principal strata—defined as response-type vectors—under multivalued treatments, multivalued outcomes, and instrumental variables (IVs). First, it systematically characterizes the identification region for such effects, equivalently reformulating it as the existence problem of solutions to a linear system subject to nonnegativity and structural constraints. Leveraging IV exogeneity and the zero-probability assumption on certain response types, the paper proposes a unified inferential framework grounded in linear programming and convex optimization—extending Fang et al. (2023). It further develops a computationally tractable, conservative, and consistent procedure for constructing confidence sets applicable to canonical causal parameters, including the population stratification effect (PSE) and variants of the complier average causal effect (CACE). The approach substantially improves inference precision and broadens applicability across complex multivalued settings.
This study addresses the limitation of conventional instrumental variable (IV) methods, which often assume a constant treatment effect and thus struggle to accommodate effect heterogeneity in real-world settings. Building on the local average treatment effect (LATE) framework, the paper systematically integrates covariate-adjusted IV approaches, clarifying how covariates influence the weighting structure of LATE estimators. It proposes flexible modeling strategies to avoid parametric misspecification and incorporates robust diagnostic tests for violations of the monotonicity assumption. By combining nonparametric and semiparametric estimation techniques, formal hypothesis testing, and accompanying software implementation, this work offers empirical researchers a theoretically rigorous yet practically feasible causal inference workflow, substantially enhancing the reliability and applicability of IV analysis.
This study addresses the estimation bias arising from endogeneity in regression models by proposing a general and computationally efficient semiparametric projection method. The approach constructs endogenous instrumental variables by projecting and expanding the conditional mean function of the structural error onto the space of explanatory variables, thereby avoiding reliance on conventional exogenous instruments or specific parametric model forms. It is applicable to linear, nonlinear, and semiparametric settings alike. By integrating LASSO-based variable selection with asymptotic theory, the paper establishes identification conditions and asymptotic properties of the resulting estimator. Extensive simulations and empirical analyses demonstrate the method’s strong finite-sample performance, confirming its practical utility in mitigating endogeneity bias across diverse modeling contexts.
This study addresses the challenge of identifying causal effects in static panel data when the treatment is endogenous and high-dimensional nonlinear confounders are present, a setting where conventional instrumental variable (IV) methods often fail. The paper proposes Panel IV-DML, the first extension of double machine learning (DML) to a panel IV framework, which integrates flexible machine learning techniques—such as Lasso and random forests—for covariate adjustment and introduces a novel weak identification diagnostic tailored to this setting. Theoretical analysis and Monte Carlo simulations demonstrate that the estimator achieves higher precision under strong instruments and more robust inference under weak instruments. Empirical applications across three immigration studies confirm that the method replicates classic 2SLS findings while also detecting scenarios of weak identification, thereby supporting more cautious causal conclusions.
Traditional instrumental variable (IV) methods often suffer from causal bias due to weak or invalid instruments and reliance on external data. To address this, we propose a novel data-driven approach that constructs synthetic instrumental variables (SIVs) solely from observed covariates. Our method introduces the “double-tilting (DT) condition”—a newly established identification criterion that enables valid IV selection without external instruments and further determines the sign of the correlation between the endogenous variable and the structural error. By integrating DT-condition testing with heteroskedasticity-robust estimation, our framework substantially improves causal effect estimation accuracy in both simulations and empirical applications. It effectively mitigates weak instrument and instrument invalidity issues while drastically reducing dependence on exogenous instruments. This work establishes a verifiable, purely observational paradigm for addressing endogeneity, advancing causal inference methodology beyond conventional IV assumptions.
本文通过代理变量调整工具变量法中的依从性加权,解决了在存在未观察到的异质性时从局部平均处理效应估计总体平均处理效应的问题。