Neighborhood Stability in Double/Debiased Machine Learning with Dependent Data

📅 2025-11-14
📈 Citations: 0
Influential: 0
📄 PDF

career value

236K/year
🤖 AI Summary
This paper addresses the efficacy of double/debiased machine learning (DML) under weakly dependent data. We propose a novel estimation framework that eliminates the need for sample-splitting via cross-fitting. Our core innovation is the introduction and generalization of “neighborhood stability” to arbitrary metric spaces—enabling seamless application to spatiotemporal and network-structured data—thereby circumventing efficiency loss and small-sample bias induced by conventional cross-fitting. Building upon Chen et al. (2022)’s stability theory and integrating weak dependence modeling techniques, the framework accommodates diverse machine learning algorithms and rigorously establishes asymptotic normality and consistency of the DML estimator in general metric spaces. The resulting estimator exhibits enhanced robustness to local data perturbations. Overall, this work provides a scalable, implementation-friendly, and theoretically grounded tool for nonparametric causal inference under dependence.

Technology Category

Application Category

📝 Abstract
This paper studies double/debiased machine learning (DML) methods applied to weakly dependent data. We allow observations to be situated in a general metric space that accommodates spatial and network data. Existing work implements cross-fitting by excluding from the training fold observations sufficiently close to the evaluation fold. We find in simulations that this can result in exceedingly small training fold sizes, particularly with network data. We therefore seek to establish the validity of DML without cross-fitting, building on recent work by Chen et al. (2022). They study i.i.d. data and require the machine learner to satisfy a natural stability condition requiring insensitivity to data perturbations that resample a single observation. We extend these results to dependent data by strengthening stability to"neighborhood stability,"which requires insensitivity to resampling observations in any slowly growing neighborhood. We show that existing results on the stability of various machine learners can be adapted to verify neighborhood stability.
Problem

Research questions and friction points this paper is trying to address.

Extends double/debiased machine learning to weakly dependent data
Establishes validity of DML without cross-fitting for dependent observations
Introduces neighborhood stability condition for spatial and network data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Extends DML to dependent data without cross-fitting
Introduces neighborhood stability for data perturbations
Adapts existing machine learners to verify stability conditions
🔎 Similar Papers