🤖 AI Summary
This study addresses identification bias in latent variable regression coefficients arising from multi-source nonlinear measurement error—exemplified by divergent measures of occupational exposure to artificial intelligence—by proposing a partial identification approach based on curvature constraints. Assuming a linear consensus measurement function and bounding heterogeneity in the curvature of individual measurement sources relative to the slope, the method constructs closed-form identification intervals that are invariant to unknown measurement loadings. These intervals exhibit sharpness, with half-widths that are second-order small relative to the curvature bounds. The curvature bounds are estimated via split-sample instrumental variable techniques, and inference with uniform coverage is achieved by combining Imbens–Manski confidence intervals with Stoye critical values. Applied to 8.88 million person-years of U.S. community survey data, the approach yields a consensus coefficient of −0.239 across five AI exposure measures, with a partial identification half-width amounting to only 1.23% of the point estimate.
📝 Abstract
We study linear regression when the regressor is latent and observed only through multiple noisy measurements, each a smooth but possibly nonlinear function of the latent variable. The problem is acute in the measurement of occupational exposure to artificial intelligence, where competing scores yield downstream estimates that differ by a factor of eleven. A regression on any single measurement recovers a source-specific coefficient rather than the structural one. We fix the latent scale by requiring the consensus measurement function to be linear and bound the remaining curvature heterogeneity across sources relative to slope. Under this bound, the structural coefficient lies in a closed-form interval centered at a symmetric cross-source estimator. The interval is invariant to unknown source loadings, and its half-width is second order in the curvature bound and sharp to the same order. With at least four measurements, the bound is estimable from the joint distribution of the sources through a split-instrument auxiliary regression, and Imbens-Manski confidence intervals with the Stoye critical value attain uniform coverage over the curvature class, including at the point-identified boundary. The application matches six exposure measures to an American Community Survey panel of 8.88 million person-year observations for 2015 to 2024. The post-2022 employment coefficient changes sign between the language-model measures and the Webb patent-text measure, and an ex ante factor-analytic rule separates the Webb measure as a distinct construct. The five retained sources yield a loading-invariant consensus coefficient of -0.239, with a partial-identification half-width of 1.23 percent of the point estimate, or 1.88 percent at the one-sided 95 percent upper bound on the curvature. We read the application as measurement reconciliation rather than as a causal estimate of AI displacement.