Score
Designs and analyzes hypothesis tests that use reproducing kernel Hilbert space (RKHS) embeddings and kernel-based statistics to represent structured null and alternative hypotheses, including kernel two-sample and other kernelized hypothesis tests as well as structured multiple-testing procedures. Works on constructing test statistics and null-sampling or allocation policies, deriving tractable finite-sample uncertainty and robustness guarantees under RKHS representations, and optimizing power subject to structural constraints.
This paper addresses the fragmentation and weak geometric intuition in existing kernel method theory by establishing a unified functional analytic framework grounded in Hilbert space geometry. Starting from the definition of positive-definite kernels, it rigorously unifies reproducing kernel Hilbert spaces (RKHS) and Hilbert–Schmidt operators, thereby reconstructing fundamental statistical concepts—including covariance, regression, and information-theoretic measures—within a coherent geometric setting. The work innovatively embeds kernel density estimation, distributional kernel embeddings, and maximum mean discrepancy (MMD) into a single RKHS paradigm, yielding a self-consistent theory bridging statistical estimation and probabilistic representation. This framework provides geometric interpretations for Gaussian processes and kernel Bayesian inference, and establishes a rigorous mathematical foundation for future theoretical advances in kernel-based learning.
To address the computational bottleneck of two-sample testing in high-dimensional non-Euclidean spaces, this paper proposes an efficient Maximum Mean Discrepancy (MMD) test based on Random Fourier Features (RFF) and spectral regularization. The method reduces time complexity from $O(n^3)$ to $O(n^2)$. It provides the first rigorous proof that RFF-based MMD attains the minimax optimal statistical rate. We further introduce a data-adaptive strategy for selecting the spectral regularization parameter and optimizing the kernel, and construct a practical permutation testing framework. Theoretical analysis, grounded in eigenvalue decay properties of integral operators, characterizes the bias–variance trade-off. Experiments on synthetic and multiple benchmark datasets demonstrate substantial computational speedup while maintaining high statistical power—achieving only a marginal loss (<3%) in test efficacy compared to the exact MMD test, and outperforming existing accelerated kernel two-sample tests.
This paper addresses the nonparametric two-sample testing problem by proposing a likelihood ratio test based on kernel Gaussian embeddings. The method jointly embeds the two distributions into a reproducing kernel Hilbert space (RKHS) via their kernel mean and kernel covariance operators, mapping them to mutually singular Gaussian measures. A regularized likelihood ratio statistic—constructed using the relative entropy between these Gaussian embeddings—is shown to converge to zero under the null hypothesis and diverge to infinity under the alternative, enabling natural hypothesis separation. The approach integrates permutation-based calibration and spectral regularization to ensure finite-sample stability. Theoretically, the test is proven to be consistent and possesses uniform power bounds. Empirically, it significantly outperforms state-of-the-art methods—including MMD and C2ST—in high-dimensional, weak-signal regimes, unifying and extending the kernel embedding-based testing framework.
This paper establishes a unified optimal theoretical framework for three kernel-based hypothesis tests: the Maximum Mean Discrepancy (MMD) two-sample test, the Hilbert–Schmidt Independence Criterion (HSIC) independence test, and the Kernel Stein Discrepancy (KSD) goodness-of-fit test. Methodologically, it derives, for the first time under the minimax statistical setting, the optimal separation rates of all three tests simultaneously—both in the $L^2$ and kernel metric topologies—and systematically characterizes the trade-offs between statistical power and practical constraints, including computational efficiency, differential privacy, and robustness to data contamination. To bridge theory and practice, the authors propose two adaptive kernel selection strategies—kernel pooling and kernel aggregation—that jointly optimize statistical efficacy and constraint satisfaction. Theoretical analysis and empirical evaluation demonstrate that the proposed methods achieve optimal separation rates across all constraint regimes. This work provides the first unified power analysis paradigm and implementable adaptive solutions for these fundamental nonparametric kernel tests.
Addressing the dual challenges of inflated Type I error rates (loss of test-level control) and low statistical power in conditional independence testing, this paper proposes a data-efficient kernel-based testing framework. The method employs kernel ridge regression and introduces, for the first time in this setting, three principled bias-correction strategies: data splitting, auxiliary data utilization, and restriction to simplified function classes—ensuring rigorous asymptotic and finite-sample control of the significance level. Theoretically, the approach guarantees convergence of the Type I error rate to the nominal significance level while enhancing detection power for complex dependency structures. Extensive experiments on diverse synthetic and real-world datasets demonstrate that the proposed method achieves precise Type I error control and substantially outperforms state-of-the-art competitors—including KCIT and RCIT—in statistical power, with improved robustness and reliability.
This work proposes a nonparametric kernel-based approach for inference in multivariate or functional time series, addressing problems such as goodness-of-fit testing, change-point detection in marginal distributions, and independence testing. The method avoids both resampling and bandwidth selection by embedding the data into a reproducing kernel Hilbert space (RKHS) and constructing test statistics through sample splitting, projection, and self-normalization. Leveraging a novel conditioning technique, the authors establish that the resulting test statistic admits a pivotal asymptotic null distribution under strong mixing conditions and analyze its power against local alternatives. The proposed procedure achieves high finite-sample accuracy while substantially improving computational efficiency, outperforming existing resampling-based methods.
This work addresses the limited power of nonparametric two-sample tests in high-dimensional or complex distributional settings by proposing the spectrally truncated normalized Maximum Mean Discrepancy (st-nMMD). Built upon embeddings in a reproducing kernel Hilbert space, st-nMMD integrates covariance operator normalization with spectral truncation regularization to substantially enhance test power. The paper establishes, for the first time, a non-asymptotic exponential upper bound for st-nMMD under the null hypothesis, introduces an adaptive hyperparameter tuning algorithm that avoids data splitting, and provides explicit non-asymptotic quantile estimates. Empirical results demonstrate that the method maintains proper Type I error control while achieving superior statistical power and stability under the alternative hypothesis, significantly outperforming existing kernel-based two-sample tests.
This work addresses the high computational cost and poor calibration of existing kernel-based conditional independence tests, which often rely on kernel ridge regression. The authors propose a regression-agnostic kernel test that models variables in a reproducing kernel Hilbert space and assesses conditional independence by testing the marginal independence of regression residuals. This approach is compatible with arbitrary regression estimators—including tree-based models—thereby overcoming the limitations imposed by dependence on specific regression forms. Built upon a generalized Hilbert–Schmidt independence criterion framework, the method provides asymptotically valid significance levels with guaranteed consistency. Empirical evaluations demonstrate that the proposed test achieves more accurate Type I error control across diverse data-generating mechanisms while exhibiting statistical power comparable to or better than state-of-the-art methods.
This work proposes a kernel ensemble $R^2$ method to address the challenge of measuring statistical dependence in multivariate, functional, and structured data under nonlinear, tail, or oscillatory dependency scenarios. By integrating local normalization with the flexibility of reproducing kernel Hilbert spaces (RKHS), the method extends ensemble $R^2$ to general kernel settings for the first time, rigorously satisfying properties such as boundedness in [0,1] and necessary and sufficient conditions for independence and deterministic functional relationships. Leveraging k-nearest neighbor graphs and conditional mean embeddings, the authors establish consistency of the graph-based estimator and derive adaptive convergence rates with respect to the intrinsic dimensionality. Experiments on both synthetic and real-world media annotation datasets demonstrate that the proposed approach significantly outperforms state-of-the-art methods, particularly in detecting nonlinear and structured dependencies with higher statistical power.
This study addresses the lack of effective tests for independence and mean independence in weakly dependent data, such as sample paths from stationary ergodic stochastic processes. Building upon the Hilbert–Schmidt Independence Criterion, the authors develop a unified testing framework applicable to general topological spaces. Under near-epoch dependence (NED) conditions, they establish, for the first time, the consistency and asymptotic distribution theory of the test statistic under both fixed and local alternatives. By integrating kernel methods with asymptotic statistical analysis, the proposed approach substantially extends the scope of existing independence tests. Its favorable finite-sample performance is demonstrated through simulation studies on functional data.