Score
Designs and implements approximations of the Maximum Mean Discrepancy (MMD) statistic using finite random feature maps (such as random Fourier features), producing low-dimensional feature representations that approximate kernel evaluations. Builds algorithms and error analyses that reduce MMD computation to linear time and enable efficient evaluation and gradient computation through the feature-based approximation.
This paper addresses the two-sample testing problem on path space—determining whether two sets of time-series paths originate from the same stochastic process. Method: We propose the signature Maximum Mean Discrepancy (sig-MMD), a kernel-based statistic built upon the signature transform, to quantify distributional discrepancies between path measures. Contribution/Results: We establish the first systematic theoretical framework for sig-MMD, identifying that statistical power degradation—particularly elevated Type-II error under finite samples—stems from the coupling of signature truncation and kernel bandwidth selection. To mitigate this, we introduce an adaptive truncation order selection scheme and a data-driven bandwidth correction strategy. Experiments across multiple synthetic and real-world path datasets demonstrate that our method reduces misclassification rates by 35%–62% compared to baselines, achieves superior robustness over existing kernel-based two-sample tests, and provides a reproducible, high-accuracy statistical inference tool for detecting differences among complex stochastic processes.
This work addresses the inconsistent performance of existing variance estimators for the Maximum Mean Discrepancy (MMD) two-sample test under varying conditions—specifically, across the null and alternative hypotheses as well as balanced and unbalanced sample settings—and the absence of a unified framework. By leveraging the U-statistic representation and Hoeffding decomposition, the authors establish the first unified, unbiased variance estimation framework for MMD that encompasses all such hypothesis and sampling configurations. Furthermore, for the one-dimensional Laplacian kernel, they develop an exact accelerated algorithm that reduces computational complexity from O(n²) to O(n log n). The proposed method demonstrates robustness in finite samples, significantly enhancing both statistical inference accuracy and computational efficiency.
This work addresses the kernel selection and bandwidth sensitivity challenges inherent in kernel-based discrepancy measures—specifically Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernel Stein Discrepancy (KSD)—for distribution comparison, independence testing, and generative model evaluation. We propose a unified computational framework and a multi-kernel adaptive fusion estimator grounded in Hilbert space embeddings and Stein operator theory. Our method integrates V- and U-statistics, employs efficient incomplete U-statistic approximations, and incorporates a data-driven bandwidth adaptation strategy. Compared to single-kernel approaches, the proposed estimator substantially improves statistical power in small-sample and high-dimensional settings, while ensuring reproducibility and ease of hyperparameter tuning. The resulting toolkit provides a theoretically coherent and practically accessible unified implementation for all three major kernel discrepancies.
This paper addresses the general parametric estimation problem without distributional assumptions, aiming to construct estimators robust to both model misspecification and complex data dependence structures. We propose a minimum distance estimation framework based on the Maximum Mean Discrepancy (MMD), establishing— for the first time—its statistical consistency under non-i.i.d. sampling and model misspecification, and providing theoretical convergence guarantees for stochastic gradient descent optimization. Theoretically, the estimator exhibits intrinsic robustness to temporal dependence, outliers, and distributional shifts. Numerical experiments demonstrate its superior performance over classical M-estimators under data contamination and intricate time-series settings. Our core contribution lies in unifying the characterization of the MMD estimator’s generalization error and robustness limits, thereby offering a novel paradigm for reliable inference with nonstandard data.
Traditional kernel Maximum Mean Discrepancy (MMD) two-sample tests rely on permutation to determine critical thresholds, ensuring finite-sample validity but incurring an O(n²) computational cost per permutation—prohibitively expensive for large samples. This paper proposes the cross-MMD test statistic: by splitting samples to construct a U-statistic, and combining studentization with a Gaussian kernel, it yields the first kernel MMD test that requires no permutations. The method achieves asymptotic normality with a single O(n²) computation, while preserving finite-sample validity, statistical consistency, and minimax optimal detection rates under local alternatives. Theoretically and empirically, cross-MMD accelerates testing by over an order of magnitude compared to permutation-based approaches on large samples, with only a marginal loss in power, and maintains strong consistency against any fixed distributional discrepancy.
This study establishes minimax lower bounds for the estimation of Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernelized Stein Discrepancy (KSD) in general topological spaces under unbounded kernel conditions. By integrating reproducing kernel Hilbert space theory, functional analysis, and a minimax information-theoretic framework, the work rigorously proves—under mild assumptions—that the optimal convergence rate for these three classes of kernel-based discrepancy measures remains $n^{-1/2}$. This result resolves a long-standing open theoretical question and extends to the estimation of mean embeddings and centered cross-covariance operators, thereby establishing the minimax optimality of their parametric convergence rates.
This work addresses the critical dependence of Maximum Mean Discrepancy (MMD) two-sample test power on kernel selection, a challenge exacerbated by existing data-driven approaches that either overfit due to violations of the i.i.d. assumption or fail to scale to continuous kernel spaces. The paper pioneers a rigorous formulation of kernel selection as a model selection problem and introduces the Complexity-Penalized MMD (CP-MMD) criterion. By deriving a complexity penalty from uniform concentration inequalities for two-sample statistics, CP-MMD seamlessly integrates into the optimization objective, enabling direct tuning of continuous kernel parameters—such as bandwidths, polynomial features, or even deep network weights—without requiring grid search. The method maintains strict Type I error control while achieving or surpassing state-of-the-art test power across diverse experimental settings.
This work addresses the lack of convergence guarantees for Maximum Mean Discrepancy (MMD) estimation in non-convex settings. By adopting the perspective of MMD gradient flows, the authors propose a Preconditioned Gradient Descent (PGD) algorithm that performs parameter optimization in the space of probability measures. They establish, for the first time, global asymptotic convergence of PGD under non-convexity by introducing gradient domination and projected residual conditions, thereby bridging nonparametric gradient flows with parametric optimization. Experimental results demonstrate that PGD significantly outperforms standard gradient descent in both parameter estimation and composite hypothesis testing tasks, corroborating both the theoretical rigor and practical efficacy of the proposed method.
This work addresses the computational and memory bottlenecks of traditional Grassmannian kernel methods, which require constructing full Gram matrices and thus struggle with high-dimensional subspace data. To overcome these limitations, the authors propose a scalable kernel approximation framework based on random rank-one projections combined with bounded nonlinear transformations—either periodic or binary—that yield compact one-bit subspace feature representations. This approach enables continuous interpolation between the inverse Binet–Cauchy kernel and Gaussian-like kernels while effectively preserving the intrinsic geometry of subspaces. The method substantially reduces computational, memory, and storage costs. Experimental results on synthetic data and the ETH-80 classification benchmark demonstrate that the proposed technique accurately maintains Grassmannian geometric relationships with high fidelity, confirming its efficiency and practical utility.
本文证明了随机特征方法在多维目标中的谱收敛性,并通过分析不同规则性的频率分布,为椭圆边界值和特征值问题提供了解决方案。