🤖 AI Summary
This study addresses the limited practicality of partial identification arising from overly wide identification regions, which often compels researchers to impose strong assumptions to achieve point identification at the cost of credibility. The authors propose a novel approach that tightens the partially identified set in linear prediction models by incorporating auxiliary moment information—such as means—provided by data publishers, without altering the original interval-censored data. This method systematically integrates moment constraints to substantially shrink the identification region without additional assumptions and elucidates how different moment conditions shape the geometry of the identified set. Leveraging tools from convex geometry and moment-constrained optimization, the paper derives directional measures of identification value under unconditional and conditional mean restrictions. Empirical analysis on survey wage data demonstrates that even minimal moment information can dramatically recover identification power lost due to data coarsening, markedly improving inference accuracy.
📝 Abstract
Partial identification is often set aside in practice because the identification regions it delivers are too wide to be useful, pushing researchers toward strong assumptions that buy point identification at the cost of credibility. We show that a source of information already sitting in most interval-valued datasets can fix this without adding any assumption at all. When an outcome is reported only as an interval---because a data custodian bracketed, top-coded, or formally privatized it to protect respondents---the same custodian typically continues to publish accurate population aggregates of that outcome, precisely because doing so does not compromise any individual record. We develop a framework for exploiting exactly this information: restricting the set of admissible completions of the data to those consistent with a known aggregate, rather than restricting the interval itself, and characterizing the sharp identification region that results for the best linear predictor. The restrictions we study behave in strikingly different ways---some collapse the region by a full dimension, others narrow it while leaving its shape intact. We characterize the geometric effect of each restriction and derive closed-form directional measures of identifying value for the mean and conditional-mean cases. An illustration using interval-valued wages from the Current Population Survey shows that the effect is far from marginal: modest auxiliary information recovers a substantial share of the identifying power usually thought to be lost once an outcome is coarsened.