🤖 AI Summary
Existing methods for testing (conditional) independence among non-Euclidean random objects struggle to balance geometric flexibility with theoretical tractability and cannot accommodate object-valued conditioning variables. This work proposes Distance Profile Embedding (DPE), a novel approach that maps random objects in general metric spaces into a Hilbert space of square-integrable functions, thereby establishing a unified framework for both marginal and conditional independence testing. DPE is the first method capable of handling object-valued conditioning variables without requiring isometric embeddings or bijective correspondence assumptions. It further provides a closed-form asymptotic null distribution, enabling analytical p-value computation. Both theoretical analysis and empirical evaluations demonstrate that DPE achieves strong performance and practical utility across synthetic data as well as real-world applications, including gut microbiome compositions and global human mortality distributions.
📝 Abstract
Testing independence or conditional independence is fundamental to statistical inference, yet existing methods for non-Euclidean random objects often face a difficult trade-off between geometric flexibility and theoretical tractability. We introduce the Distance Profile Embedding (DPE), a novel representation that maps random objects from general metric spaces into a Hilbert space of square-integrable functions. We prove that this mapping is injective and preserves full distributional information without requiring isometric Hilbert embeddings or one-to-one correspondence conditions. Leveraging the DPE, we develop a unified framework for marginal and conditional independence testing of random objects that enjoys a rigorous asymptotic theory for both size and power. Notably, our framework is the first in the literature to accommodate object-valued conditioning variables when testing conditional independence, overcoming the Euclidean or Hilbertian constraints of existing methodologies. We facilitate the calculation of analytic $p$-values using closed-form asymptotic null distributions, which avoids the computational burden of permutation tests common in existing metric-based methods. The numerical properties of our methods are demonstrated through both simulations and two real-world applications involving gut microbiome compositions and global human mortality distributions, respectively.