๐ค AI Summary
This paper addresses the critical yet underexplored problem of testing equivalence between two conditional distributionsโa fundamental task in transfer learning and causal inference. We propose the first unified framework for both global and local two-sample conditional distribution testing. Our method introduces: (1) novel distance and kernel-based metrics that characterize conditional distribution homogeneity; (2) an estimation theory grounded in conditional U-statistics, enabling integrated modeling for both global and local tests; and (3) a principled combination of RKHS embeddings and localized bootstrap resampling, yielding convergence rates and asymptotic null/alternative distributions of the estimators. Theoretical analysis guarantees strong statistical power, while empirical evaluations on synthetic and real-world datasets demonstrate high detection accuracy and robustness.
๐ Abstract
Testing the equality of two conditional distributions is crucial in various modern applications, including transfer learning and causal inference. Despite its importance, this fundamental problem has received surprisingly little attention in the literature. This work aims to present a unified framework based on distance and kernel methods for both global and local two-sample conditional distribution testing. To this end, we introduce distance and kernel-based measures that characterize the homogeneity of two conditional distributions. Drawing from the concept of conditional U-statistics, we propose consistent estimators for these measures. Theoretically, we derive the convergence rates and the asymptotic distributions of the estimators under both the null and alternative hypotheses. Utilizing these measures, along with a local bootstrap approach, we develop global and local tests that can detect discrepancies between two conditional distributions at global and local levels, respectively. Our tests demonstrate reliable performance through simulations and real data analyses.