On the use of cross-fitting in causal machine learning with correlated units

📅 2026-01-15
📈 Citations: 0
Influential: 0
📄 PDF

career value

250K/year
🤖 AI Summary
In causal machine learning, it remains unclear whether standard cross-fitting can effectively eliminate bias introduced by black-box algorithms when observational units exhibit spatial, clustered, or time-series dependencies. This study systematically evaluates the performance of conventional cross-fitting that ignores such dependence structures through theoretical analysis and simulation experiments. The findings reveal that, even without explicitly modeling inter-unit dependencies, standard cross-fitting successfully removes the dominant bias term and yields estimation bias and precision comparable to—or sometimes better than—specialized decorrelated folding strategies across a range of correlated data-generating mechanisms. These results challenge the prevailing assumption in the literature that customized cross-fitting procedures are necessary for dependent data, offering theoretical justification for simplifying causal inference pipelines.

Technology Category

Application Category

📝 Abstract
In causal machine learning, the fitting and evaluation of nuisance models are typically performed on separate partitions, or folds, of the observed data. This technique, called cross-fitting, eliminates bias introduced by the use of black-box predictive algorithms. When study units may be correlated, such as in spatial, clustered, or time-series data, investigators often design bespoke forms of cross-fitting to minimize correlation between folds. We prove that, perhaps contrary to popular belief, this is typically unnecessary: performing cross-fitting as if study units were independent usually still eliminates key bias terms even when units may be correlated. In simulation experiments with various correlation structures, we show that causal machine learning estimators typically have the same or improved bias and precision under cross-fitting that ignores correlation compared to techniques striving to eliminate correlation between folds.
Problem

Research questions and friction points this paper is trying to address.

causal machine learning
cross-fitting
correlated units
nuisance models
bias reduction
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-fitting
causal machine learning
correlated data
nuisance models
bias reduction
🔎 Similar Papers
No similar papers found.
S
Salvador V. Balkus
Department of Biostatistics, Harvard Chan School of Public Health
H
Hasan Laith
Department of Statistics, Harvard College
N
N. Hejazi
Department of Biostatistics, Harvard Chan School of Public Health