🤖 AI Summary
Accurately estimating the direct, unmediated effect of environmental amenities on housing prices (DUET) is crucial for welfare analysis, yet existing methodologies face significant limitations. This study leverages over one million property transactions in New York State from 1990 to 2024 to construct an empirical “ground truth” that preserves the true data-generating process. Using Monte Carlo simulations, it systematically evaluates the estimation accuracy of generalized difference-in-differences (DID), two-way fixed effects models, and causal machine learning methods—including causal forest DID—across varying sample sizes. The results demonstrate that generalized DID consistently outperforms benchmark models across all scenarios, while causal machine learning approaches exhibit strong performance when sample sizes exceed 3,000 observations, with causal forest DID achieving accuracy nearly on par with generalized DID, thereby offering a reliable methodological alternative for DUET estimation.
📝 Abstract
Hedonic price models are widely used to assess how environmental amenities affect property values, yet methodological guidance for estimating direct price effects remains sparse. We conduct an empirical Monte Carlo simulation to evaluate the performance of traditional and causal machine learning approaches for estimating the direct unmediated price effect of spatially delineated amenities on treated properties (DUET), a conservative lower-bound approximation for welfare changes with direct applications to benefit-cost analysis. Where previous simulations rely on parametric assumptions, we retain the actual data-generating process underlying over 1 million property transactions from upstate New York (1990--2024). By randomly assigning "treatment locations" across iterations we establish a "ground truth" that allows us to precisely measure estimation error. Our results demonstrate that generalized difference-in-differences (DID) regression consistently outperforms baseline DID and two-way fixed effects models across all scenarios. Causal Machine Learning (CML) methods, particularly causal forest DID, achieve comparable performance to generalized DID in most scenarios. In larger samples (above 3,000 treated) increasingly common in contemporary hedonic studies, CML approaches offer substantial advantages when properly specified. Based on empirical simulation results, we provide a set of method-specific best practice recommendations for both traditional regression and causal machine learning approaches.