Inference after data-driven control-unit selection in difference-in-differences with estimated covariance

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inferential bias in the average treatment effect on the treated (ATT) arising from data-driven control group selection and covariance matrix estimation within difference-in-differences frameworks. To mitigate this, it proposes a selective inference framework that extends exact Gaussian procedures to settings with unknown covariance by employing plug-in estimators to correct for selection bias, applicable to both individual panel and repeated cross-sectional data. This work unifies inference under complex scenarios involving staggered adoption, treatment effect heterogeneity, and sample reuse. Furthermore, it establishes uniform conditional and marginal coverage guarantees, proves the asymptotic equivalence between plug-in estimators and known-covariance confidence intervals, and derives convergence rates for interval lengths, thereby ensuring the statistical validity of ATT confidence intervals.
📝 Abstract
In difference-in-differences (DiD), researchers may use pre-treatment trends to select a control group for which the parallel-trends assumption appears plausible, with the aim of estimating the average treatment effect on the treated (ATT). Our earlier paper,Nakano and Hoshino (2016), and the present paper jointly provide the first selective-inference approach to the ATT that explicitly accounts for this control selection. We generalize our exact Gaussian procedure with known covariance to allow the covariance matrix to be estimated from the same individual-level data used for control selection and DiD estimation. We use this estimate to compute the variance, conditioning direction, residual, and truncation set. With fixed numbers of regions and periods, we establish uniform conditional coverage for selection events with probabilities bounded away from zero, and marginal coverage of the selected target without that restriction. We allow unequal regional sample sizes, heterogeneous covariances, ties in population fit, and regional sample shares that converge to zero. We establish asymptotic equivalence between the plug-in and known-covariance interval endpoints and derive rates for interval length. For staggered adoption, the control pools may differ across cohorts and periods, controls may be not yet treated, observations may be reused, and treatment effects may be heterogeneous. We also construct inference conditional on unions of selection paths that leave the reported parameter unchanged, together with simultaneous confidence bands for finitely many event-time effects. Under parallel trends and the other identifying conditions, the coverage results apply to the ATT. We give sufficient sampling conditions for individual panels and independent repeated cross-sections.
Problem

Research questions and friction points this paper is trying to address.

Difference-in-Differences
Selective Inference
Data-driven Control Selection
Estimated Covariance
Average Treatment Effect on the Treated
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective Inference
Difference-in-Differences
Estimated Covariance
Staggered Adoption
Average Treatment Effect on the Treated
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
R
Ryoya Nakano
Graduate School of Economics, Keio University, 2-15-45 Mita, Minato-ku, Tokyo 108-8345, Japan
T
Takahiro Hoshino
Faculty of Economics, Keio University, 2-15-45 Mita, Minato-ku, Tokyo 108-8345, Japan; RIKEN Center for Advanced Intelligence Project, 1-4-1 Nihonbashi, Chuo-ku, Tokyo 103-0027, Japan