Econometric Inference with Machine-Learned Proxies: Partial Identification via Data Combination

📅 2026-04-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the pervasive issue of estimation bias and invalid inference that arises when machine learning–generated proxy variables are directly employed in downstream econometric models. The authors propose a novel identification framework that leverages two data sources: a downstream sample containing covariates and the proxy, and an auxiliary validation sample comprising the proxy alongside its true target variable. Using the proxy as a bridge, they construct a locally identified model based on unconditional optimal transport. Crucially, this approach does not require the upstream machine learning estimator to be consistent or to satisfy specific convergence rates, nor does it rely on a fully observed validation sample. Valid asymptotic inference with correct size is achieved through analytically derived critical values, eliminating the need for resampling. Monte Carlo simulations demonstrate that the method maintains accurate size control and yields informative confidence sets across a range of proxy prediction accuracies.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationIntelligent Robots: State EstimationReasoning under Uncertainty: Stochastic Optimization

Application Category

Economics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAIUser Modeling, Personalization and Recommendation: User privacy protection in personalized systemsResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Empirical researchers increasingly use upstream machine-learning (ML) methods to construct proxies for latent target variables from complex, unstructured data. A naive plug-in use of such proxies in downstream econometric models, however, can lead to biased estimation and invalid inference. This paper develops a framework for partial identification and inference in general moment models with ML-generated proxies. Our approach does not require restrictive assumptions on the upstream ML procedure, such as consistency or known convergence rates, nor does it require a complete validation sample containing all variables used in the downstream analysis. Instead, we assume access to two datasets: a downstream sample containing observed covariates and the proxy, and an auxiliary validation sample containing joint observations on the proxy and its target variable. We treat the proxy as a linking variable between these two samples, rather than as a literal noisy substitute for the latent target variable. Building on this idea, we develop a sharp identification strategy based on an unconditional optimal transport characterization and an inference procedure that controls asymptotic size using analytical critical values without resampling. Monte Carlo simulations show reliable size control and informative confidence sets across a range of predictive-accuracy scenarios.
Problem

Research questions and friction points this paper is trying to address.

partial identification
machine learning proxies
econometric inference
data combination
moment models
Innovation

Methods, ideas, or system contributions that make the work stand out.

partial identification
machine-learned proxies
optimal transport
data combination
econometric inference
L
Lixiong Li
Johns Hopkins University