🤖 AI Summary
This study addresses the issue that accumulated estimation errors in generative conditional independence testing often lead to uncontrolled Type I error. To mitigate this, we reformulate conditional generative modeling as a multi-source domain adaptation problem and propose DA-Diff, a multi-source domain adaptation diffusion model, along with the DA-CIT testing framework. By integrating conditional diffusion models with weighted empirical risk minimization, this framework leverages multi-source auxiliary data to enhance the accuracy of target-domain distribution estimation. Theoretically, we establish the controllability of asymptotic Type I error. Empirically, our approach significantly improves conditional generation quality, achieving strict Type I error control while maintaining competitive statistical power.
📝 Abstract
Conditional independence (CI) is a fundamental concept in statistics and machine learning. Recent advances in conditional generative modeling provide flexible tools for generative-model-based CI tests, which rely on an estimated conditional distribution to generate randomized samples. However, errors in estimating this distribution accumulate in existing Type I error bounds, and consistency of the generative estimator alone does not guarantee asymptotic Type I error control. To address this limitation, we formulate conditional generative modeling as a domain adaptation problem and leverage auxiliary data from multiple source domains to improve estimation in the target CI testing domain. We propose Domain-Adapted Diffusion (DA-Diff), a multi-source domain adaptation framework for conditional diffusion models based on weighted empirical risk minimization over both target and source domains. We establish the convergence rate of DA-Diff and show how transferable source data can improve target-domain estimation through an increased effective sample size while controlling transfer bias. Building on DA-Diff, we further propose Domain-Adapted Conditional Independence Testing (DA-CIT) and show that its Type I error satisfies $P(p \leq \alpha) \leq \alpha + o(1)$. Experiments demonstrate that DA-Diff improved conditional generation quality compared with transfer-learning diffusion baselines, while DA-CIT provides strong Type I error control and competitive power.