Score
Design and implement regularization methods based on the Maximum Mean Discrepancy (MMD) for unit‑norm (spherical) representations, including deterministic computation of full‑dimensional discrepancy and analytic integration over random projection directions. Use these methods to enforce or measure uniformity of distributions on the hypersphere and to reduce gradient variance introduced by projection‑based MMD estimators.
To address the limited discriminative power of Maximum Mean Discrepancy (MMD) in goodness-of-fit testing, this paper proposes a spectral-filtering-based regularized kernel discrepancy framework. The method constructs flexible test statistics via integral operators, relaxing stringent assumptions on kernels and filter functions inherent in prior approaches, and achieves rigorous Type-I error control with improved statistical power in non-asymptotic settings. Theoretically, the proposed test constitutes a natural generalization of existing MMD-based tests, offering enhanced detection sensitivity and broader theoretical applicability. Empirical evaluations demonstrate that it matches or outperforms state-of-the-art methods across diverse scenarios—including multivariate, high-dimensional, and small-sample settings—exhibiting strong practical adaptability and competitive performance.
This work addresses the inconsistent performance of existing variance estimators for the Maximum Mean Discrepancy (MMD) two-sample test under varying conditions—specifically, across the null and alternative hypotheses as well as balanced and unbalanced sample settings—and the absence of a unified framework. By leveraging the U-statistic representation and Hoeffding decomposition, the authors establish the first unified, unbiased variance estimation framework for MMD that encompasses all such hypothesis and sampling configurations. Furthermore, for the one-dimensional Laplacian kernel, they develop an exact accelerated algorithm that reduces computational complexity from O(n²) to O(n log n). The proposed method demonstrates robustness in finite samples, significantly enhancing both statistical inference accuracy and computational efficiency.
Existing gradient flow methods for source-to-target distribution transport face a fundamental trade-off: f-divergence-based flows lack numerical tractability, while MMD-based flows require strong assumptions—such as explicit noise injection—to ensure convergence. This work proposes DrMMD, a tractable and robust gradient flow method that operates solely on target samples. Its core innovation is the first-established tunable de-regularized linkage between MMD and the χ²-divergence, unifying near-global convergence guarantees with closed-form sample update rules. DrMMD integrates de-regularized kernel MMD, the Wasserstein gradient flow framework, and an adaptive scheduling strategy, ensuring theoretical convergence for general target distributions in both continuous- and discrete-time settings. Extensive experiments on large-scale teacher–student neural networks validate its effectiveness, robustness, and scalability.
This work addresses local structure modeling of point clouds in product spaces endowed with mixed Euclidean and directional metrics. We propose the first subspace-constrained mean shift algorithm tailored to such hybrid metric spaces, enabling joint estimation of density modes and density ridges. Theoretically, we establish convergence guarantees on product manifolds and provide practical implementation criteria. By integrating manifold gradient analysis with a customized product-space metric design, the method enhances interpretability and fidelity in capturing heterogeneous multi-source structures. Experiments on synthetic and real-world data—including 3D human poses and motion trajectories—demonstrate substantial improvements over state-of-the-art approaches in mode and ridge localization accuracy, robustness to noise, and structural interpretability. Our framework establishes a new paradigm for density-based topological modeling in complex geometric domains.
Kernelized Stein discrepancy (KSD) suffers from theoretical limitations in controlling weak convergence and precisely separating target distributions. Method: We integrate Bochner embedding theory, Stein’s method, kernel analysis, and weak topology theory to systematically address these limitations. Contribution/Results: First, we establish the necessary and sufficient conditions for KSD to metrize weak convergence—its first rigorous characterization. Second, we construct a novel class of unbounded kernels that are universally discriminative—capable of separating all Borel probability measures—overcoming the inherent discriminability constraints of bounded kernels. Third, we propose the first KSD variant provably equivalent to weak convergence. Our framework significantly enhances KSD’s separation power and convergence control: on ℝᵈ, it enables precise quantitative characterization of weak convergence toward any target distribution P. This advancement strengthens theoretical guarantees and empirical performance in statistical hypothesis testing, sample quality assessment, and Stein variational gradient descent (SVGD) sampling.
This work addresses the instability in existing self-supervised learning methods that rely on slice-based regularization via random one-dimensional projections, which introduces high gradient variance. The authors propose a full-dimensional statistical regularization objective directly defined on the unit hypersphere, eliminating the need for stochastic projection approximations. They establish, for the first time, an analytical equivalence between slice-based regularization and Maximum Mean Discrepancy (MMD), and introduce deterministic regularizers based on MMD, Kernelized Stein Discrepancy (KSD), and KL divergence. Leveraging spectral theory, they construct rotation-invariant kernels—Heat and Bandlimited—to enable unbiased and stable distribution matching over the hypersphere. Experiments on ImageNet and Galaxy10 demonstrate faster convergence, improved training stability, and superior performance. Notably, different statistical criteria induce distinct representation geometries, with KL-based regularization achieving the best results in texture retrieval tasks.
This study establishes minimax lower bounds for the estimation of Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernelized Stein Discrepancy (KSD) in general topological spaces under unbounded kernel conditions. By integrating reproducing kernel Hilbert space theory, functional analysis, and a minimax information-theoretic framework, the work rigorously proves—under mild assumptions—that the optimal convergence rate for these three classes of kernel-based discrepancy measures remains $n^{-1/2}$. This result resolves a long-standing open theoretical question and extends to the estimation of mean embeddings and centered cross-covariance operators, thereby establishing the minimax optimality of their parametric convergence rates.
This work addresses the lack of convergence guarantees for Maximum Mean Discrepancy (MMD) estimation in non-convex settings. By adopting the perspective of MMD gradient flows, the authors propose a Preconditioned Gradient Descent (PGD) algorithm that performs parameter optimization in the space of probability measures. They establish, for the first time, global asymptotic convergence of PGD under non-convexity by introducing gradient domination and projected residual conditions, thereby bridging nonparametric gradient flows with parametric optimization. Experimental results demonstrate that PGD significantly outperforms standard gradient descent in both parameter estimation and composite hypothesis testing tasks, corroborating both the theoretical rigor and practical efficacy of the proposed method.
This work addresses the convergence challenges of Maximum Mean Discrepancy (MMD) gradient flows arising from non-convexity by proposing a Sobolev-regularized MMD gradient flow method with gradient penalty on the witness function. The approach eliminates the need for isoperimetric assumptions on the target distribution and, for the first time, guarantees global convergence simultaneously in both continuous and discrete time. It provides a unified framework applicable to sampling and generative modeling tasks involving unnormalized densities. By integrating Sobolev regularization, kernel mean embeddings, and Stein kernel techniques, the proposed method demonstrates superior convergence properties and generalization performance over existing approaches across a range of experiments.
Existing approaches, such as those based on the von Mises–Fisher (vMF) distribution, model only the mean direction and thus fail to capture complex geometric structures—such as multimodality, axial symmetry, or zonal patterns—in spherical weighted empirical measures. This work proposes a Geometric Information Decomposition (GID) framework that leverages spherical harmonics to construct a nested sequence of maximum-entropy projections, hierarchically quantifying the incremental information-theoretic gaps at each level. For the first time, this enables a layered decomposition of higher-order geometric structures inherent in spherical measures, transcending the limitations of single-parameter models by fully characterizing features ranging from the mean direction to high-order anisotropy and fine angular patterns. Theoretical guarantees include invariance, consistency, and asymptotic normality, along with a quadratic-form zero-calibration test. Experiments on circular and spherical data successfully reveal latent structures invisible to vMF-based methods, demonstrating the approach’s efficacy and practical utility.