🤖 AI Summary
This paper addresses the challenge of identifying heterogeneous treatment effects in regression discontinuity designs (RDD) without prior specification of effect-modifying covariates. We propose an “honest” RDD tree method grounded in supervised machine learning, which—novelty—incorporates honest splitting to ensure valid statistical inference while automatically selecting pre-treatment covariates that drive heterogeneity, without requiring prespecified candidate variables. By integrating RDD theory, decision-tree modeling, and Monte Carlo simulation, our approach achieves superior bias control and nominal coverage of confidence intervals compared to conventional subgroup analyses or interaction-based methods. An empirical application to Romanian secondary education data successfully uncovers multiple sources of treatment-effect heterogeneity, demonstrating the method’s robustness and practical utility for causal inference in RDD settings.
📝 Abstract
The paper proposes a supervised machine learning algorithm to uncover treatment effect heterogeneity in classical regression discontinuity (RD) designs. Extending Athey and Imbens (2016), I develop a criterion for building an honest"regression discontinuity tree", where each leaf of the tree contains the RD estimate of a treatment (assigned by a common cutoff rule) conditional on the values of some pre-treatment covariates. It is a priori unknown which covariates are relevant for capturing treatment effect heterogeneity, and it is the task of the algorithm to discover them, without invalidating inference. I study the performance of the method through Monte Carlo simulations and apply it to the data set compiled by Pop-Eleches and Urquiola (2013) to uncover various sources of heterogeneity in the impact of attending a better secondary school in Romania.