🤖 AI Summary
This work addresses the challenge of conditional independence testing in high-dimensional, small-sample settings where heterogeneous external or unlabeled data are available. Existing methods suffer from inflated Type I error rates and reduced power due to the difficulty of accurately estimating unknown conditional distributions. To overcome this, the authors propose the CRT* framework, which innovatively integrates a smoothed residual bootstrap with transfer learning to consistently estimate the conditional distribution. CRT* further enhances power by adaptively fusing information through an optimal convex combination of test statistics. Theoretical analysis establishes the consistency of the conditional distribution estimator and demonstrates, for the first time, that heterogeneous data can be effectively leveraged while strictly controlling Type I error. Empirical results on both simulated data and real-world breast cancer RNA-seq datasets show that CRT* substantially outperforms standard CRT in statistical power.
📝 Abstract
The conditional randomization test (CRT) provides a principled approach to conditional independence (CI) testing, guaranteeing exact type-I error control when the true conditional distribution is known. In practice, however, this distribution must be estimated, and estimation errors can inflate type-I errors, while high dimensionality and limited sample sizes can reduce power. Although external and unlabeled data offer the potential to improve CI testing, naive integration that ignores distributional heterogeneity can compromise type-I error control and fail to enhance power. We propose \textbf{CRT*}, a novel framework that robustly integrates external and unlabeled datasets to enhance CI testing in heterogeneous scenarios. CRT* employs smooth residual-bootstrap (SRB) with transfer learning for conditional distribution estimation, combined with adaptive data fusion via an optimal convex combination of test statistics. We theoretically establish that the SRB-based estimator converges to the true conditional distribution in expected total variation distance. Furthermore, even in high-dimensional regimes, CRT* maintains valid type-I error control and achieves strictly higher power than standard CRT without external data. Simulations and RNA-seq breast cancer data analyses demonstrate that CRT* substantially improves power while maintaining type-I error control in heterogeneous settings.