🤖 AI Summary
This study addresses the problem of insufficient confidence interval coverage in cross-fitting under non-regular settings, where cross-fold correlation undermines inferential validity. By establishing a central limit theorem grounded in locality conditions, this work reveals the asymptotic normality of cross-fitted estimators despite non-regularity. Furthermore, it introduces a novel method for estimating cross-fold correlation to correct the asymptotic variance, integrating random forests and neural networks to facilitate valid statistical inference. The primary contribution lies in theoretically quantifying and adjusting for the impact of cross-fold dependence. Simulation experiments demonstrate that the proposed approach achieves near-nominal coverage rates across diverse models, substantially enhancing the reliability of machine learning-based inference in non-regular scenarios.
📝 Abstract
Cross-fitting is routine in much of applied research. While conventional confidence intervals that ignore cross-fold dependence are asymptotically valid in several settings, they undercover in many applications that share a common form of nonregularity: from the classic cross-validation problem of testing whether a fitted model outperforms another, to testing for heterogeneous treatment effects with machine learning, to estimating the value of a potentially non-unique optimal treatment regime. Exploiting a new locality condition, I show that a large class of cross-fitting estimators still satisfies a central limit theorem despite the nonregularity, but with an asymptotic variance that must be adjusted for the cross-fold correlation. Then, I propose a method for estimating this correlation and construct new confidence intervals that attain asymptotically nominal coverage. Finally, I show that the proposed confidence intervals attain approximately nominal coverage in a simulation study with random forests and neural networks.