selection instrument

Designs and evaluates selection instruments — variables (instrumental variables) intended to affect only the selection mechanism or treatment assignment and not the outcome directly. Builds diagnostics and estimation procedures (e.g., for Heckman-style corrections) to restore identification and valid coverage in the presence of selection bias, and to detect when such corrections are unidentifiable.

selectioninstrument

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Learning control variables and instruments for causal analysis in observational data

Jul 05, 2024
NA
Nicolas Apfel
🏛️ University of Innsbruck | University of York | University of Fribourg | Heinrich Heine University Düsseldorf

Estimating causal effects from observational data requires selecting appropriate control and instrumental variables that satisfy causal identification conditions—a challenging task often reliant on strong domain knowledge or ad hoc assumptions. Method: This paper proposes the first end-to-end joint learning framework that automatically identifies valid combinations of control and instrumental variables. Grounded in conditional independence testing, the method integrates nonparametric dependence measures with structural search optimization, ensuring statistical consistency in variable selection under mild regularity conditions. Contribution/Results: Unlike conventional approaches requiring prespecified variable sets or strong prior assumptions, our framework is fully data-driven. In simulations, it achieves significantly higher variable identification accuracy. Empirically, applied to the Job Corps study, its estimated treatment effect closely aligns with results from the randomized controlled trial—demonstrating both validity and robustness in real-world causal inference.

Detects control variables and instruments for causal analysis in observational dataLearns partition of instruments and control variables from observed dataTests joint existence of instruments and control variables using machine learning

Existing instrumental variable (IV) methods for nonignorable missingness (MNAR) impose restrictive a priori bounds on selection bias in the outcome scale, leading to underestimated uncertainty. This paper proposes a novel IV framework that achieves nonparametric identification of the missing outcome distribution under a multiplicative selection model and a no-interaction assumption—yielding point identification without constraining the magnitude of selection bias, a theoretical first. The method constructs a semiparametric, multiply robust estimator based on influence functions, accommodating both discrete (multivalued) and continuous IVs, and is applicable to complex survey data. Simulation studies demonstrate strong finite-sample performance. We apply the method to HIV survey data from Botswana, using interviewer characteristics as instruments to correct for dependence-induced nonresponse bias.

Develops a multiplicative instrumental variable model for MNAR dataIdentifies missing outcome functionals without restricting selection biasProvides robust estimators for nonparametric models with continuous instruments

Constructing an Instrument as a Function of Covariates

Mar 13, 2025
MS
Moses Stewart
🏛️ Harvard

This paper exposes the severe nonrobustness of instrument variables (IVs) constructed via nonlinear transformations of covariates under mild model misspecification. When exogenous IVs are unavailable, researchers often generate instruments from functional transformations of observed covariates; however, we theoretically demonstrate that—even under a constant linear treatment effect—any modest nonlinear misspecification in the true structural function induces arbitrarily large bias in the resulting IV estimator. This is the first rigorous theoretical characterization of the extreme sensitivity of such constructed IVs to nonlinearities in the structural function, challenging the widely adopted empirical practice of “safe construction.” Combining asymptotic theory with semi-synthetic experiments—calibrating real data to multiple structural models—we empirically confirm substantial deviations of IV estimates from the true causal effect. Our findings provide a critical robustness warning for IV construction in applied econometrics and causal inference.

Assesses robustness of IV specifications to structural nonlinearityExamines bias in IV estimand with covariate-constructed instrumentsInvestigates reliability of IV estimates under misspecification

Policy-relevant causal effect estimation using instrumental variables with interference

Sep 15, 2025
DN
Didier Nibbering
🏛️ Monash University | University of Lisbon

Conventional instrumental variable (IV) methods assume no interference among units, yet real-world policy evaluations frequently involve social interactions that violate this assumption. Method: Under a mild interference assumption, we formally define policy-interpretable direct and spillover effects, and achieve partial identification of these effects under generalized monotonic treatment response and selection assumptions—without imposing parametric restrictions on the interference structure. Our approach integrates IV identification strategies, the potential outcomes framework, and multi-peer interference modeling to derive computationally tractable bounds on causal effects. Contribution/Results: We break the no-interference barrier and provide, for the first time, policy-relevant bounds on causal effects for IV designs with social interactions. This significantly extends the applicability and credibility of IV methods in evaluating real-world social programs—such as education and public health interventions—where peer effects are prevalent.

Estimating causal effects with interference in IV methodsExtending IV estimation to realistic social contextsIdentifying direct and spillover effects under mild assumptions

Revisiting the Many Instruments Problem using Random Matrix Theory

Aug 16, 2024
HF
Helmut Farbmacher
🏛️ Technical University of Munich

Instrumental variable (IV) estimation suffers from severe finite-sample bias when the number of instruments $p$ far exceeds the sample size $n$. This paper systematically introduces random matrix theory to high-dimensional IV settings, revealing the implicit bias–variance trade-off advantage of ridge regularization under dense first-stage regressions—and extending this analysis to the $p > n$ regime. By reconstructing the finite-sample bias structure of two-stage least squares (2SLS), we propose a unified correction framework grounded in random matrix asymptotics, substantially improving second-stage estimation accuracy. We establish theoretical consistency of the proposed estimator under both high-dimensional sparse and dense first-stage designs. Empirically, the method reduces estimation error by over 30% on average across benchmark specifications. Our approach unifies and generalizes existing bias approximation and correction theories for high-dimensional IV estimation.

Addressing bias in instrumental variables with many instrumentsConnecting traditional bias adjustments to Silverstein equationGeneralizing asymptotic properties to high-dimensional instrument settings

Latest Papers

What's happening recently
View more

This study addresses causal effect identification in observational settings with unmeasured confounding and potentially invalid instrumental variables, focusing on linear instrumental variable models with multiple endogenous treatments. The authors propose generalized majority and plurality rules to achieve identification, coupled with a data-driven instrument selection procedure that yields sampling confidence intervals robust to the erroneous inclusion of invalid instruments. Under standard regularity conditions, these intervals are shown to attain asymptotic nominal coverage and exhibit length shrinking at the parametric rate. The practical utility and validity of the proposed method are demonstrated through an empirical application in Mendelian randomization.

causal inferenceinstrumental variablesinvalid instruments

Traditional instrumental variable methods rely on the stringent assumption that the structural equation model holds exactly—a condition often violated in practice, leading to invalid inference. This work proposes a novel inference framework based on debiased least squares and inverse problem regularization, which defines a target parameter that coincides with conventional estimands when the structural model is correctly specified yet remains well-defined and inferable even under model misspecification. By relaxing the requirement of exact structural equation validity, the approach ensures robust statistical inference under substantially weaker conditions, thereby significantly enhancing the reliability and applicability of instrumental variable methods in realistic settings.

Debiased InferenceInstrumental VariableInverse Problems

This study addresses the challenge of efficiently estimating causal effects under confounding when experimental budgets are limited. The authors propose a novel approach that integrates instrumental variable regression with Gaussian graphical models, leveraging prior knowledge of partial joint distributions to optimize the allocation between fully observed samples and partially observed data (e.g., only \(X_{12}\)). Under a fixed budget constraint, this method analytically derives the optimal sampling scheme that minimizes the asymptotic variance of the causal effect estimator—a solution not previously available in closed form. Theoretical analysis demonstrates that the proposed allocation significantly reduces both the total budget and the number of complete observations required to detect non-zero causal effects. Empirical validation in automotive analytics and drug discovery underscores the method’s practical utility alongside its theoretical contributions.

budget constraintcausal effect estimationexperimental design

This study addresses the quantification of omitted variable bias in nonlinear instrumental variable (IV) estimation by extending sensitivity analysis to nonlinear IV frameworks, encompassing local average treatment effects (LATE), LATE for treated individuals (LATT), and partially linear IV models (PLIVM). The authors derive bias decompositions, construct partial identification bounds, and develop computable bias bounds alongside robust inference procedures that adjust confidence intervals accordingly. Integrating double machine learning (DML), the approach accommodates flexible control for high-dimensional covariates. Application to the JTPA experiment reveals that estimated program effects for women remain robustly significant, whereas those for men are sensitive to potential omitted variables; first-stage compliance rate estimates are stable, but intent-to-treat and treatment effect estimates exhibit greater fragility.

nonlinear instrumental variableomitted variable biaspartial identification

This study addresses unmeasured confounding in case–control studies arising from time-invariant direct effects of instrumental variables, thereby relaxing the conventional exclusion restriction. The authors propose an instrumental variable–based difference-in-differences approach tailored for retrospective designs. Built upon structural mean models and incorporating corrections for case–control sampling and selection bias, this method uniquely integrates instrumental variable analysis with the difference-in-differences framework while permitting time-invariant direct effects of the instrument on the outcome. Simulation studies demonstrate favorable finite-sample performance, and the approach is successfully applied to a French health insurance database to effectively assess the risk of serious infections associated with biologic therapy for psoriasis.

case-control designdifference-in-differencesexclusion restriction

Hot Scholars

FB

Francois Buet-Golfouse

Barclays
Financial MathematicsMachine LearningFairnessMulti-Objective Reinforcement Learning
ZS

Zekai Shao

Fudan University
VisualizationHuman Computer InteractionHuman-AI CollaborationVisual Analytics
RW

Rui Wang

Professor of Population Medicine (Biostatistics), Harvard University and Harvard Pilgrim
biostatisticsclinical trialsHIVcancer
LS

Leixian Shen

The Hong Kong University of Science and Technology
Human-AI CollaborationVisual Data AnalysisData Storytelling
AJ

Ayush Jha

PhD Candidate: Economics, Department of Economics, Texas Tech University
FinanceTime Series EconometricsMarket Microstructure