Score
Design, implement, and analyze kernel-based two-sample hypothesis tests that use the maximum mean discrepancy (MMD) in a reproducing kernel Hilbert space (RKHS) to decide equality of two distributions, including construction of test statistics and selection or combination of kernels to achieve omnibus sensitivity to differences in location, scale, shape, and tails. Derive the test's asymptotic behavior under the null (often a Gaussian quadratic-form limit), provide finite-sample calibration procedures (e.g., permutation or bootstrap), and characterize power and limiting distributions under local alternatives.
Traditional kernel Maximum Mean Discrepancy (MMD) two-sample tests rely on permutation to determine critical thresholds, ensuring finite-sample validity but incurring an O(n²) computational cost per permutation—prohibitively expensive for large samples. This paper proposes the cross-MMD test statistic: by splitting samples to construct a U-statistic, and combining studentization with a Gaussian kernel, it yields the first kernel MMD test that requires no permutations. The method achieves asymptotic normality with a single O(n²) computation, while preserving finite-sample validity, statistical consistency, and minimax optimal detection rates under local alternatives. Theoretically and empirically, cross-MMD accelerates testing by over an order of magnitude compared to permutation-based approaches on large samples, with only a marginal loss in power, and maintains strong consistency against any fixed distributional discrepancy.
To address the high computational cost, low statistical power, and bandwidth sensitivity of kernel two-sample tests on high-dimensional, large-scale data, this paper proposes a parameter-free robust kernel test. The method avoids bandwidth selection entirely while ensuring reliability and high power. Its core contributions are threefold: (1) a novel test statistic designed via theoretical analysis to eliminate power loss from data splitting; (2) a non-asymptotic significance control mechanism guaranteeing validity under finite samples; and (3) inherent suitability for high-dimensional settings, delivering uniformly high power across diverse alternative hypotheses. Experiments on synthetic and real-world datasets demonstrate that the proposed method achieves 10–100× speedup over MMD and state-of-the-art large-scale kernel tests, with average power gains of 15%–40%, all without any bandwidth tuning.
This paper addresses the multivariate two-sample and $k$-sample goodness-of-fit testing problem by introducing the first unified kernelized quadratic distance (KQD) framework. Methodologically, it integrates both two-sample and $k$-sample tests within a single statistical model grounded in matrix-valued distances and reproducing kernel Hilbert space (RKHS) theory; derives the asymptotic null distribution of the test statistic rigorously; and enables finite-sample inference via Monte Carlo or permutation procedures. Key contributions include: (i) establishing the first theoretical equivalence between KQD-based tests and maximum mean discrepancy (MMD) tests; (ii) proposing a scalable, statistically rigorous paradigm for multi-group testing; and (iii) releasing QuadratiK, an open-source software package supporting both R and Python. Extensive simulations and real-data analyses demonstrate that the method maintains accurate Type-I error control while achieving substantial gains in statistical power.
This paper establishes a unified optimal theoretical framework for three kernel-based hypothesis tests: the Maximum Mean Discrepancy (MMD) two-sample test, the Hilbert–Schmidt Independence Criterion (HSIC) independence test, and the Kernel Stein Discrepancy (KSD) goodness-of-fit test. Methodologically, it derives, for the first time under the minimax statistical setting, the optimal separation rates of all three tests simultaneously—both in the $L^2$ and kernel metric topologies—and systematically characterizes the trade-offs between statistical power and practical constraints, including computational efficiency, differential privacy, and robustness to data contamination. To bridge theory and practice, the authors propose two adaptive kernel selection strategies—kernel pooling and kernel aggregation—that jointly optimize statistical efficacy and constraint satisfaction. Theoretical analysis and empirical evaluation demonstrate that the proposed methods achieve optimal separation rates across all constraint regimes. This work provides the first unified power analysis paradigm and implementable adaptive solutions for these fundamental nonparametric kernel tests.
To address the $O(n^2)$ computational bottleneck of kernelized Stein discrepancy (KSD) under large-scale data—arising from its reliance on U- or V-statistics—this paper introduces, for the first time, the Nyström low-rank kernel approximation into KSD estimation, yielding a scalable and accelerated KSD estimator. The proposed method reduces time complexity to $O(mn + m^3)$, where $m ll n$, and establishes $sqrt{n}$-consistency under sub-Gaussian assumptions. Theoretical analysis is grounded in the Stein operator and reproducing kernel Hilbert space (RKHS) framework, balancing statistical efficiency with computational tractability. Extensive benchmark experiments demonstrate that the new estimator retains statistical power comparable to the original KSD while substantially enhancing practicality for large-scale goodness-of-fit testing. This work provides an efficient, theoretically sound tool for high-dimensional distribution fitting and hypothesis testing.
This work addresses the limited power of nonparametric two-sample tests in high-dimensional or complex distributional settings by proposing the spectrally truncated normalized Maximum Mean Discrepancy (st-nMMD). Built upon embeddings in a reproducing kernel Hilbert space, st-nMMD integrates covariance operator normalization with spectral truncation regularization to substantially enhance test power. The paper establishes, for the first time, a non-asymptotic exponential upper bound for st-nMMD under the null hypothesis, introduces an adaptive hyperparameter tuning algorithm that avoids data splitting, and provides explicit non-asymptotic quantile estimates. Empirical results demonstrate that the method maintains proper Type I error control while achieving superior statistical power and stability under the alternative hypothesis, significantly outperforming existing kernel-based two-sample tests.
This study establishes minimax lower bounds for the estimation of Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernelized Stein Discrepancy (KSD) in general topological spaces under unbounded kernel conditions. By integrating reproducing kernel Hilbert space theory, functional analysis, and a minimax information-theoretic framework, the work rigorously proves—under mild assumptions—that the optimal convergence rate for these three classes of kernel-based discrepancy measures remains $n^{-1/2}$. This result resolves a long-standing open theoretical question and extends to the estimation of mean embeddings and centered cross-covariance operators, thereby establishing the minimax optimality of their parametric convergence rates.
This work addresses the need for more efficient, robust, and flexible metrics for measuring distances between probability distributions in statistical inference and numerical integration. Centered on kernel methods, we propose an efficient estimator for Maximum Mean Discrepancy (MMD), develop novel MMD-based approaches for conditional expectation estimation and integral calibration, and introduce a new family of distance measures—kernel quantile discrepancies—that effectively overcome MMD’s limitations in tail sensitivity and discriminative power. Both theoretical analysis and empirical experiments demonstrate that the proposed methods offer strong scalability, computational efficiency, and superior performance, thereby providing more powerful and practical kernel-based tools for nonparametric statistics and integration tasks.
This work addresses the inconsistent performance of existing variance estimators for the Maximum Mean Discrepancy (MMD) two-sample test under varying conditions—specifically, across the null and alternative hypotheses as well as balanced and unbalanced sample settings—and the absence of a unified framework. By leveraging the U-statistic representation and Hoeffding decomposition, the authors establish the first unified, unbiased variance estimation framework for MMD that encompasses all such hypothesis and sampling configurations. Furthermore, for the one-dimensional Laplacian kernel, they develop an exact accelerated algorithm that reduces computational complexity from O(n²) to O(n log n). The proposed method demonstrates robustness in finite samples, significantly enhancing both statistical inference accuracy and computational efficiency.
This work addresses the critical dependence of Maximum Mean Discrepancy (MMD) two-sample test power on kernel selection, a challenge exacerbated by existing data-driven approaches that either overfit due to violations of the i.i.d. assumption or fail to scale to continuous kernel spaces. The paper pioneers a rigorous formulation of kernel selection as a model selection problem and introduces the Complexity-Penalized MMD (CP-MMD) criterion. By deriving a complexity penalty from uniform concentration inequalities for two-sample statistics, CP-MMD seamlessly integrates into the optimization objective, enabling direct tuning of continuous kernel parameters—such as bandwidths, polynomial features, or even deep network weights—without requiring grid search. The method maintains strict Type I error control while achieving or surpassing state-of-the-art test power across diverse experimental settings.