Score
Designs and analyzes statistical hypothesis tests that assess whether the conditional mean of an outcome is independent of covariates by constructing test statistics based on graph or nearest-neighbor relationships. This work typically involves studentizing nearest-neighbor estimators, deriving asymptotic standard-normal null distributions, and constructing asymptotically level-correct tests without relying on bootstrap or sample splitting.
This work addresses a key limitation in existing nonparametric tests for significance and conditional independence, which often rely on the strong assumption that covariate densities have compact support, thereby restricting their applicability. To overcome this, the authors propose a weighted function approach based on nonparametric orthogonal projections, coupled with a computationally efficient multiplier bootstrap procedure to accurately estimate critical values. The framework is further extended to conditional independence testing. By relaxing the compact support requirement, the method achieves parametric-rate detection power against local alternatives under more general density conditions. Finite-sample experiments demonstrate that the proposed approach substantially improves both the accuracy and stability of hypothesis testing compared to existing methods.
This study addresses the challenge of node-level two-sample hypothesis testing in Gaussian graphical models, where existing methods struggle with precise localization on decomposable graphs and exhibit instability under small sample sizes. The authors propose a leave-one-out Bartlett-corrected likelihood ratio test based on fully connected graphs, which enables calibrated significance inference for individual nodes and fixed-size node subsets—a capability not previously achieved. By integrating a leave-one-out strategy with Bartlett correction, the method constructs a test statistic whose null distribution asymptotically follows a standard chi-squared distribution, thereby overcoming limitations inherent in traditional clique-based decomposition approaches. Simulations demonstrate that the proposed test achieves excellent calibration and statistical power, and its practical utility is further corroborated through real-data analysis.
In graphical model structure learning, numerous redundant conditional independence (CI) tests—unused by mainstream algorithms—are performed; while some can detect or correct model errors, not all possess such error-correcting capability. Method: This paper establishes the first systematic semantic taxonomy of redundant CI tests, grounded in graphical model theory, the axiomatic system of conditional independence, and analysis of probabilistic representability. It rigorously identifies that only CI statements entailed by the graph structure—not by universal probabilistic properties—exhibit genuine error-correction potential. Contribution/Results: We formalize a semantic classification framework for redundant CI tests, precisely delineating their applicability boundaries and principled conditions for error detection and model correction. This work enhances the robustness of structure learning and provides a theoretical foundation for designing fault-tolerant graphical model discovery algorithms.
This paper addresses the critical yet underexplored problem of testing equivalence between two conditional distributions—a fundamental task in transfer learning and causal inference. We propose the first unified framework for both global and local two-sample conditional distribution testing. Our method introduces: (1) novel distance and kernel-based metrics that characterize conditional distribution homogeneity; (2) an estimation theory grounded in conditional U-statistics, enabling integrated modeling for both global and local tests; and (3) a principled combination of RKHS embeddings and localized bootstrap resampling, yielding convergence rates and asymptotic null/alternative distributions of the estimators. Theoretical analysis guarantees strong statistical power, while empirical evaluations on synthetic and real-world datasets demonstrate high detection accuracy and robustness.
Addressing the dual challenges of inflated Type I error rates (loss of test-level control) and low statistical power in conditional independence testing, this paper proposes a data-efficient kernel-based testing framework. The method employs kernel ridge regression and introduces, for the first time in this setting, three principled bias-correction strategies: data splitting, auxiliary data utilization, and restriction to simplified function classes—ensuring rigorous asymptotic and finite-sample control of the significance level. Theoretically, the approach guarantees convergence of the Type I error rate to the nominal significance level while enhancing detection power for complex dependency structures. Extensive experiments on diverse synthetic and real-world datasets demonstrate that the proposed method achieves precise Type I error control and substantially outperforms state-of-the-art competitors—including KCIT and RCIT—in statistical power, with improved robustness and reliability.
This study addresses the fundamental question of whether a statistical parameter defined through a conditional distribution remains constant across covariates—a problem encompassing treatment effect heterogeneity and conditional association. The authors propose a general nonparametric testing framework based on smooth functionals applied to conditional distributions, yielding functional parameters for which they construct test statistics with tractable asymptotic distributions. Their approach explicitly links to norm-based tests in function spaces and, compared to existing norm-type methods, exhibits superior asymptotic properties under the null hypothesis. Simulation studies demonstrate strong finite-sample performance, and the method is successfully applied to data from a breast cancer clinical trial, effectively identifying key biomarkers predictive of response to adjuvant chemotherapy.
This work proposes a nearest-neighbor-graph-based estimator for the proportion of explained variation by the conditional mean function, offering a unified framework for model evaluation, model-free variable screening, and higher-order Sobol’ interaction analysis. The proposed estimator achieves nearly linear computational complexity and enables asymptotically standard normal inference without resorting to bootstrapping or sample splitting, thereby guaranteeing correct significance levels and universal consistency. Theoretical results establish its convergence rate and asymptotic normality, while extensive simulations and real-data analyses demonstrate substantial improvements over existing methods in both computational efficiency and statistical power.
Existing nonparametric conditional independence tests lack a unified, efficient, and theoretically guaranteed approach for mixed-type data involving both discrete and continuous variables. This work proposes a graph-based kernel method that constructs composite neighborhoods by exact matching on discrete variables and k-nearest neighbors on continuous variables, augmented with local polynomial debiasing to correct smoothing bias. The method establishes, for the first time, asymptotic null distribution theory across all combinations of variable types, achieving a dimension-agnostic detection rate of \(n^{-1/4}\). It circumvents the phase transition issues inherent in high-dimensional geometric estimators and reduces computational complexity from cubic to nearly quadratic. The proposed test thus offers statistical consistency, computational scalability, and rigorous theoretical guarantees.
This study addresses the absence of a statistical inference theory for the Azadkia–Chatterjee conditional dependence coefficient estimator \( T_n \) under general dependence structures. We establish its asymptotic normality and provide, for the first time, a central limit theorem valid for arbitrary dependence. By integrating rank-based and nearest-neighbor constructions, we derive a closed-form expression for the asymptotic variance and propose a consistent variance estimator computable in \( O(n \log n) \) time. Combined with existing bias-correction techniques, our results yield a complete inferential framework for this measure, substantially broadening its applicability in practical data analysis.
This study addresses the challenges of assumption validity, robustness, and scalability in conditional independence (CI) testing for constraint-based causal discovery with high-dimensional mixed-type data. It systematically reviews and compares six major classes of CI methods—partial correlation, contingency tables, regression residuals, k-nearest neighbors, kernel-based approaches, and machine learning techniques—evaluating their performance and failure modes under diverse data-generating mechanisms, small sample sizes, and heterogeneous variable types. For the first time, it comprehensively delineates the applicability boundaries of each method and elucidates how their errors propagate into inaccuracies in causal skeleton and v-structure identification. The work also surveys current implementations in R and Python libraries and identifies key future directions, including discretization-free CI tests for mixed data, improved error control in small samples, and enhanced scalability.