Score
Designs and implements causal-forest estimators adapted to difference-in-differences setups to flexibly estimate heterogeneous treatment effects across units and time, including algorithm adaptations, sample-splitting, tuning, and prediction of conditional average treatment effects. Builds accompanying identification checks, inference procedures, and evaluation workflows that exploit DID variation to recover treatment-effect heterogeneity under realistic data-generating processes and compare performance to generalized DID approaches.
This paper addresses the failure of conventional difference-in-differences (DiD) estimators under staggered treatment adoption and heterogeneous treatment effects. We propose DiD-BCF—a novel framework that reparameterizes the parallel trends assumption as an identifiable nonlinear function, thereby relaxing standard linearity and homogeneity restrictions. Integrating Bayesian Causal Forests (BCF), nonparametric modeling, and the DiD structure, DiD-BCF employs MCMC-based posterior inference and tree ensembles to deliver unified, robust estimation of the average treatment effect (ATE), group-level ATE (GATE), and conditional ATE (CATE). Simulations demonstrate substantial gains over leading DiD methods in settings with nonlinearity, selection bias, and strong heterogeneity. Empirically, applying DiD-BCF to U.S. minimum wage policy reveals significant population-size–dependent conditional treatment effects—heterogeneity entirely overlooked by conventional DiD approaches—highlighting the model’s improved identification power and substantive interpretability.
该论文提出一种基于决策树和随机森林的算法,通过结合两种分裂标准来预测和解释个体治疗效果的异质性和偏差问题。
In clinical practice, relative risk (RR) is more clinically interpretable than absolute risk difference; however, mainstream heterogeneous treatment effect (HTE) methods—such as causal forests—rely on absolute risk differences for recursive partitioning, often overlooking RR heterogeneity. To address this, we propose the first RR-oriented variant of causal forest: a nonparametric node-splitting criterion grounded in generalized linear model comparisons, explicitly designed to maximize statistical power for detecting RR heterogeneity. Our method imposes no strong distributional assumptions and automatically identifies covariates—and their interaction structures—that drive RR heterogeneity. Simulation studies and empirical analyses demonstrate that the proposed framework substantially improves identification of clinically relevant subgroups, successfully capturing RR heterogeneity patterns missed by conventional causal forests. It achieves this while preserving computational feasibility, thereby enhancing both clinical interpretability and statistical power.
Standard Causal Forests applied to panel data often yield spurious heterogeneity in treatment effects due to their neglect of unit and time fixed effects. This work proposes a node-level residualization strategy that locally removes fixed effects within each candidate split during tree construction, thereby avoiding the bias introduced by global demeaning. Building upon the Causal Forests framework, the method integrates fixed effects modeling with random forests to efficiently estimate heterogeneous treatment effects. The authors implement and publicly release the approach as the Python package `causalfe`. Simulation studies demonstrate that the proposed method accurately recovers true treatment effect heterogeneity across a range of data-generating processes and substantially outperforms conventional Causal Forests.
This study addresses the problem of extrapolating causal effects from multi-site randomized controlled trials (RCTs) to a new target site with baseline survey data only. To handle site-level population heterogeneity and unobserved confounding, we propose modeling baseline covariates as functional data—thereby capturing site-specific confounding structures—for the first time. We then develop a design-oriented, nonparametric method to construct an optimal finite-dimensional feature space, ensuring optimal convergence rates for conditional average treatment effect (CATE) estimation. Our approach integrates functional data analysis, nonparametric regression, and causal transfer learning theory. Evaluated across five integrated multi-site RCTs on cash transfer programs, the method significantly improves prediction accuracy of treatment effects at target sites and quantifies the estimation gain attributable to adaptive transfer.
This study addresses a critical limitation of fixed-effects causal forests in estimating conditional average treatment effects (CATE): the systematic attenuation of effect heterogeneity due to averaging within leaf nodes, which underestimates the true variability of treatment effects. The authors demonstrate that this bias stems from the leaf-node averaging mechanism and formally characterize its dependence on signal-to-noise ratio, panel size, and dimensionality. To correct this attenuation, they propose a self-contained, out-of-bag cross-fitting debiasing procedure that recovers the true CATE distribution by rescaling the estimated slopes. The method integrates causal forests, within-transformation for fixed effects, and optimal linear prediction, and is implemented in the `causalfe` Python package. Simulations show that the proposed correction reduces CATE mean squared error by 25–42% compared to standard recentering, and an empirical application to county-level minimum wage data successfully restores previously compressed effect heterogeneity.
This study addresses the challenges of estimating heterogeneous causal effects and discovering subgroups under imperfect compliance by proposing a two-step, model-agnostic framework. Methodologically, it introduces a non-Bayesian tree structure compatible with arbitrary machine learning algorithms, integrating DRRF-IV, GRF-IV, and forest learner techniques to achieve precise estimation of the compiler average causal effect (CCACE). Experimental results demonstrate that the proposed method attains accuracy comparable to BCF-IV while significantly improving computational efficiency. Furthermore, when applied to clinical ICU data, the framework successfully identifies specific beneficiary subgroups, providing a reliable basis for individualized medical decision-making.
This study addresses the bias in conventional difference-in-differences estimators under staggered policy adoption, which arises from treatment effect heterogeneity and leads to “forbidden comparisons.” To overcome this issue, the authors propose a fixed-effects causal forest method that residualizes both outcomes and treatment indicators within each group-time unit using unit and time fixed effects. By integrating an honest causal tree splitting mechanism, the approach directly eliminates confounding through within-node differencing without requiring explicit modeling of nuisance functions, thereby enabling nonparametric estimation of the covariate-conditional group-time average treatment effect $\tau_{g,t}(x)$. Monte Carlo simulations demonstrate that the estimator is unbiased and yields valid confidence interval coverage under staggered adoption and cohort-level heterogeneity. Applied to Medicaid expansion data, the method estimates an average 2.25-percentage-point reduction in uninsurance rates, with substantially larger effects in low-income counties.
本文通过结合因果生存森林与负控制方法,提出了一种新的非参数异质性治疗效应估计方法NC-CSF,以解决观察性生存研究中因删失结果和未测量混杂因素引起的问题。
该研究解决了因果森林中因变量编码冗余导致的预测不一致问题,通过分组生成相同分割的变量来恢复预测不变性。