Better Measurement or Larger Samples? Data Collection for Policy Learning with Unobserved Heterogeneity

πŸ“… 2026-04-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF

career value

200K/year
πŸ€– AI Summary
This study addresses the challenge of optimizing data collection to enhance social welfare in policy learning when unobserved heterogeneity is present. Accounting for latent individual differences in policy responses, the authors propose a repeated-measurement design based on proxy variables for latent traits and derive minimax regret bounds for policy rules that either incorporate or omit these latent variables. The theoretical analysis uncovers a novel trade-off between policy class complexity and estimation accuracy, leading to an optimal data collection strategy that allocates resources efficiently between measurement precision and sample size. In a development economics application, incorporating a proxy for entrepreneurs’ managerial ability increases social welfare by 5% and reduces the probability of welfare loss by 50%.

Technology Category

Application Category

πŸ“ Abstract
Empirical research shows that individuals' responses to treatments vary along latent characteristics, such as innate ability or motivation. Therefore, a policymaker seeking to maximize welfare may consider designing policies based on observed characteristics and estimated latent traits. I characterize how the estimates' precision affects the worst-case performance of policies, deriving rate-sharp regret bounds for assignment rules that include or exclude them, highlighting new trade-offs with the policy space complexity. I then study how a policymaker can solve such trade-offs by designing tailored data collections, and derive the minimax optimal collection plan. In an empirical application in development economics, I show that including a proxy for entrepreneurs' business skills in targeting cash transfers increases welfare by 5%, and halves the probability of generating welfare losses. Moreover, I estimate the optimal allocation of resources between improving the precision of the proxy via repeated measurements and increasing sample size.
Problem

Research questions and friction points this paper is trying to address.

policy learning
unobserved heterogeneity
data collection
treatment effects
welfare maximization
Innovation

Methods, ideas, or system contributions that make the work stand out.

unobserved heterogeneity
policy learning
rate-sharp regret bounds
minimax optimal data collection
latent traits
πŸ”Ž Similar Papers