Score
Designs and constructs instrumental variables derived from the model itself—e.g., projections of residuals or functions of endogenous variables and controls—to enable identification and estimation without relying on external instruments, applicable in parametric and nonparametric settings and with attention to computational tractability.
This paper addresses the challenge that the “rich-covariates” condition—critical for identification in instrumental variable (IV) estimation—is often violated when instruments are non-randomly assigned and the model is non-saturated. To resolve this, we propose two nonparametric correction strategies: constructing a deconditioned instrument or augmenting the regressor set with nonparametrically estimated conditional expectations (e.g., via kernel or series regression). Our approach is the first to systematically correct for confounding bias in non-saturated, non-randomized experimental settings without relying on model saturation or instrument randomness—as assumed in Blandhol et al. (2025). We establish consistency and asymptotic normality of the resulting IV estimator under mild regularity conditions. Monte Carlo simulations demonstrate that, in finite samples, our estimator exhibits substantially lower bias, improved confidence interval coverage, and greater robustness compared to conventional IV methods.
This study addresses the estimation bias arising from endogeneity in regression models by proposing a general and computationally efficient semiparametric projection method. The approach constructs endogenous instrumental variables by projecting and expanding the conditional mean function of the structural error onto the space of explanatory variables, thereby avoiding reliance on conventional exogenous instruments or specific parametric model forms. It is applicable to linear, nonlinear, and semiparametric settings alike. By integrating LASSO-based variable selection with asymptotic theory, the paper establishes identification conditions and asymptotic properties of the resulting estimator. Extensive simulations and empirical analyses demonstrate the method’s strong finite-sample performance, confirming its practical utility in mitigating endogeneity bias across diverse modeling contexts.
Traditional instrumental variable methods rely on the stringent assumption that the structural equation model holds exactly—a condition often violated in practice, leading to invalid inference. This work proposes a novel inference framework based on debiased least squares and inverse problem regularization, which defines a target parameter that coincides with conventional estimands when the structural model is correctly specified yet remains well-defined and inferable even under model misspecification. By relaxing the requirement of exact structural equation validity, the approach ensures robust statistical inference under substantially weaker conditions, thereby significantly enhancing the reliability and applicability of instrumental variable methods in realistic settings.
Estimating causal effects from observational data requires selecting appropriate control and instrumental variables that satisfy causal identification conditions—a challenging task often reliant on strong domain knowledge or ad hoc assumptions. Method: This paper proposes the first end-to-end joint learning framework that automatically identifies valid combinations of control and instrumental variables. Grounded in conditional independence testing, the method integrates nonparametric dependence measures with structural search optimization, ensuring statistical consistency in variable selection under mild regularity conditions. Contribution/Results: Unlike conventional approaches requiring prespecified variable sets or strong prior assumptions, our framework is fully data-driven. In simulations, it achieves significantly higher variable identification accuracy. Empirically, applied to the Job Corps study, its estimated treatment effect closely aligns with results from the randomized controlled trial—demonstrating both validity and robustness in real-world causal inference.
Instrumental variable (IV) estimation suffers from severe finite-sample bias when the number of instruments $p$ far exceeds the sample size $n$. This paper systematically introduces random matrix theory to high-dimensional IV settings, revealing the implicit bias–variance trade-off advantage of ridge regularization under dense first-stage regressions—and extending this analysis to the $p > n$ regime. By reconstructing the finite-sample bias structure of two-stage least squares (2SLS), we propose a unified correction framework grounded in random matrix asymptotics, substantially improving second-stage estimation accuracy. We establish theoretical consistency of the proposed estimator under both high-dimensional sparse and dense first-stage designs. Empirically, the method reduces estimation error by over 30% on average across benchmark specifications. Our approach unifies and generalizes existing bias approximation and correction theories for high-dimensional IV estimation.
This study addresses nonparametric point identification in multivariate instrumental variable models with continuous endogenous variables when only binary instruments are available. The authors generalize the rank invariance assumption from univariate settings to cyclic monotonicity in the first stage and construct multivariate ranks via the inverse Brenier map, thereby overcoming limitations of conventional approaches that rely on restrictive heterogeneity structures or large instrument support. This framework substantially broadens the class of identifiable distributions, accommodating non-quasiconcave and multimodal densities. Moreover, the paper establishes a verifiable nonparametric identification condition that holds generically across common parametric distribution families and is empirically testable under mild non-degeneracy assumptions.
This paper addresses the nonparametric identification of demand for differentiated products, relaxing the conventional assumption that product characteristics are exogenous. Its key innovation is a novel “credibility” condition that enables nonparametric identification of counterfactual prices using only exogenous supply-side instruments—such as cost shocks or policy interventions—without requiring product attributes to serve as instruments. Methodologically, it reformulates the demand system via recentered instrumental variables and verifies identifiability under joint completeness and credibility conditions. Theoretically, it extends the nonparametric instrumental variable (IV) framework by establishing the general validity of this approach under diverse, non-nested pricing mechanisms and broad generalized index models. This significantly enhances both the credibility and applicability of demand estimation in realistic market settings.
This study addresses nonparametric inference on treatment effects in separable binary treatment selection models under unobserved confounding. By introducing a perturbation-independent parametrization directly defined by observed data and combining instrumental variables with a fixed-point argument, the authors construct a semiparametrically efficient estimator that avoids imposing unnecessary restrictions on the nuisance functions. The approach flexibly accommodates nonlinear effects, average treatment effects over the full population, and non-randomly missing data, while allowing modern machine learning methods to estimate nuisance components. Theoretical analysis establishes efficiency bounds and validates the underlying generative model, and both simulations and an empirical application to the Job Corps data demonstrate that the method yields efficient and robust estimation of smooth functionals, along with testable identification assumptions.
This study addresses the challenge of efficiently estimating causal effects under confounding when experimental budgets are limited. The authors propose a novel approach that integrates instrumental variable regression with Gaussian graphical models, leveraging prior knowledge of partial joint distributions to optimize the allocation between fully observed samples and partially observed data (e.g., only \(X_{12}\)). Under a fixed budget constraint, this method analytically derives the optimal sampling scheme that minimizes the asymptotic variance of the causal effect estimator—a solution not previously available in closed form. Theoretical analysis demonstrates that the proposed allocation significantly reduces both the total budget and the number of complete observations required to detect non-zero causal effects. Empirical validation in automotive analytics and drug discovery underscores the method’s practical utility alongside its theoretical contributions.
This work proposes a novel framework based on conditional conformal inference to construct prediction intervals for nonparametric instrumental variable (NPIV) regression that are distribution-free and offer finite-sample coverage guarantees—features absent in existing methods. By reformulating the conditional coverage problem as marginal coverage over a user-specified class of instrumental variable perturbations, the approach yields valid prediction intervals compatible with any NPIV estimator, including sieve two-stage least squares and neural network-based minimax estimators. The method provides, for the first time, distribution-free finite-sample coverage guarantees for NPIV regression while allowing practitioners to tailor the perturbation class to reflect domain-specific assumptions, thereby balancing flexibility with rigorous theoretical validity.