spherical mmd regularization

Design and implement regularization methods based on the Maximum Mean Discrepancy (MMD) for unit‑norm (spherical) representations, including deterministic computation of full‑dimensional discrepancy and analytic integration over random projection directions. Use these methods to enforce or measure uniformity of distributions on the hypersphere and to reduce gradient variance introduced by projection‑based MMD estimators.

sphericalmmdregularization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Integral-Operator-Based Spectral Algorithms for Goodness-of-Fit Tests

Nov 10, 2025
SS
Shiwei Sang
🏛️ Xi'an Jiaotong University

To address the limited discriminative power of Maximum Mean Discrepancy (MMD) in goodness-of-fit testing, this paper proposes a spectral-filtering-based regularized kernel discrepancy framework. The method constructs flexible test statistics via integral operators, relaxing stringent assumptions on kernels and filter functions inherent in prior approaches, and achieves rigorous Type-I error control with improved statistical power in non-asymptotic settings. Theoretically, the proposed test constitutes a natural generalization of existing MMD-based tests, offering enhanced detection sensitivity and broader theoretical applicability. Empirical evaluations demonstrate that it matches or outperforms state-of-the-art methods across diverse scenarios—including multivariate, high-dimensional, and small-sample settings—exhibiting strong practical adaptability and competitive performance.

Developing goodness-of-fit tests with valid error control and enhanced powerGeneralizing kernel-based discrepancy measures using spectral filtering techniquesImproving MMD's ability to distinguish between distributions through regularization

This work addresses the inconsistent performance of existing variance estimators for the Maximum Mean Discrepancy (MMD) two-sample test under varying conditions—specifically, across the null and alternative hypotheses as well as balanced and unbalanced sample settings—and the absence of a unified framework. By leveraging the U-statistic representation and Hoeffding decomposition, the authors establish the first unified, unbiased variance estimation framework for MMD that encompasses all such hypothesis and sampling configurations. Furthermore, for the one-dimensional Laplacian kernel, they develop an exact accelerated algorithm that reduces computational complexity from O(n²) to O(n log n). The proposed method demonstrates robustness in finite samples, significantly enhancing both statistical inference accuracy and computational efficiency.

Imbalanced DataMaximum Mean DiscrepancyTwo-sample Testing

(De)-regularized Maximum Mean Discrepancy Gradient Flow

Sep 23, 2024
ZC
Zonghao Chen
🏛️ University College London | Pennsylvania State University | Gatsby Computational Neuroscience Unit | ENSAE | CREST | Institut Polytechnique de Paris

Existing gradient flow methods for source-to-target distribution transport face a fundamental trade-off: f-divergence-based flows lack numerical tractability, while MMD-based flows require strong assumptions—such as explicit noise injection—to ensure convergence. This work proposes DrMMD, a tractable and robust gradient flow method that operates solely on target samples. Its core innovation is the first-established tunable de-regularized linkage between MMD and the χ²-divergence, unifying near-global convergence guarantees with closed-form sample update rules. DrMMD integrates de-regularized kernel MMD, the Wasserstein gradient flow framework, and an adaptive scheduling strategy, ensuring theoretical convergence for general target distributions in both continuous- and discrete-time settings. Extensive experiments on large-scale teacher–student neural networks validate its effectiveness, robustness, and scalability.

Develops gradient flow for transporting source to target distributionsEnsures convergence for broad target classes with sample-based implementationUses adaptive de-regularization to balance discretization errors and divergence

This work addresses local structure modeling of point clouds in product spaces endowed with mixed Euclidean and directional metrics. We propose the first subspace-constrained mean shift algorithm tailored to such hybrid metric spaces, enabling joint estimation of density modes and density ridges. Theoretically, we establish convergence guarantees on product manifolds and provide practical implementation criteria. By integrating manifold gradient analysis with a customized product-space metric design, the method enhances interpretability and fidelity in capturing heterogeneous multi-source structures. Experiments on synthetic and real-world data—including 3D human poses and motion trajectories—demonstrate substantial improvements over state-of-the-art approaches in mode and ridge localization accuracy, robustness to noise, and structural interpretability. Our framework establishes a new paradigm for density-based topological modeling in complex geometric domains.

Estimating local modes in Euclidean and directional product spacesExtending mean shift algorithm to handle product space challengesValidating convergence and effectiveness on simulated and real data

Targeted Separation and Convergence with Kernel Discrepancies

Sep 26, 2022
AB
A. Barp
🏛️ University College London | The Alan Turing Institute | Mirelo AI | University of Cambridge | Microsoft Research

Kernelized Stein discrepancy (KSD) suffers from theoretical limitations in controlling weak convergence and precisely separating target distributions. Method: We integrate Bochner embedding theory, Stein’s method, kernel analysis, and weak topology theory to systematically address these limitations. Contribution/Results: First, we establish the necessary and sufficient conditions for KSD to metrize weak convergence—its first rigorous characterization. Second, we construct a novel class of unbounded kernels that are universally discriminative—capable of separating all Borel probability measures—overcoming the inherent discriminability constraints of bounded kernels. Third, we propose the first KSD variant provably equivalent to weak convergence. Our framework significantly enhances KSD’s separation power and convergence control: on ℝᵈ, it enables precise quantitative characterization of weak convergence toward any target distribution P. This advancement strengthens theoretical guarantees and empirical performance in statistical hypothesis testing, sample quality assessment, and Stein variational gradient descent (SVGD) sampling.

Characterize kernels for separating Bochner embeddable measuresDevelop conditions for weak convergence control with bounded kernelsExpand KSD conditions to metrize convergence for hypothesis testing

Latest Papers

What's happening recently
View more

This work addresses the instability in existing self-supervised learning methods that rely on slice-based regularization via random one-dimensional projections, which introduces high gradient variance. The authors propose a full-dimensional statistical regularization objective directly defined on the unit hypersphere, eliminating the need for stochastic projection approximations. They establish, for the first time, an analytical equivalence between slice-based regularization and Maximum Mean Discrepancy (MMD), and introduce deterministic regularizers based on MMD, Kernelized Stein Discrepancy (KSD), and KL divergence. Leveraging spectral theory, they construct rotation-invariant kernels—Heat and Bandlimited—to enable unbiased and stable distribution matching over the hypersphere. Experiments on ImageNet and Galaxy10 demonstrate faster convergence, improved training stability, and superior performance. Notably, different statistical criteria induce distinct representation geometries, with KL-based regularization achieving the best results in texture retrieval tasks.

HypersphereProjection VarianceRepresentation Collapse

This study establishes minimax lower bounds for the estimation of Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernelized Stein Discrepancy (KSD) in general topological spaces under unbounded kernel conditions. By integrating reproducing kernel Hilbert space theory, functional analysis, and a minimax information-theoretic framework, the work rigorously proves—under mild assumptions—that the optimal convergence rate for these three classes of kernel-based discrepancy measures remains $n^{-1/2}$. This result resolves a long-standing open theoretical question and extends to the estimation of mean embeddings and centered cross-covariance operators, thereby establishing the minimax optimality of their parametric convergence rates.

HSICkernel discrepancyKSD

This work addresses the lack of convergence guarantees for Maximum Mean Discrepancy (MMD) estimation in non-convex settings. By adopting the perspective of MMD gradient flows, the authors propose a Preconditioned Gradient Descent (PGD) algorithm that performs parameter optimization in the space of probability measures. They establish, for the first time, global asymptotic convergence of PGD under non-convexity by introducing gradient domination and projected residual conditions, thereby bridging nonparametric gradient flows with parametric optimization. Experimental results demonstrate that PGD significantly outperforms standard gradient descent in both parameter estimation and composite hypothesis testing tasks, corroborating both the theoretical rigor and practical efficacy of the proposed method.

Minimum MMD estimationnon-convexityoptimization problem

This work addresses the convergence challenges of Maximum Mean Discrepancy (MMD) gradient flows arising from non-convexity by proposing a Sobolev-regularized MMD gradient flow method with gradient penalty on the witness function. The approach eliminates the need for isoperimetric assumptions on the target distribution and, for the first time, guarantees global convergence simultaneously in both continuous and discrete time. It provides a unified framework applicable to sampling and generative modeling tasks involving unnormalized densities. By integrating Sobolev regularization, kernel mean embeddings, and Stein kernel techniques, the proposed method demonstrates superior convergence properties and generalization performance over existing approaches across a range of experiments.

global convergencegradient flowMaximum Mean Discrepancy

Existing approaches, such as those based on the von Mises–Fisher (vMF) distribution, model only the mean direction and thus fail to capture complex geometric structures—such as multimodality, axial symmetry, or zonal patterns—in spherical weighted empirical measures. This work proposes a Geometric Information Decomposition (GID) framework that leverages spherical harmonics to construct a nested sequence of maximum-entropy projections, hierarchically quantifying the incremental information-theoretic gaps at each level. For the first time, this enables a layered decomposition of higher-order geometric structures inherent in spherical measures, transcending the limitations of single-parameter models by fully characterizing features ranging from the mean direction to high-order anisotropy and fine angular patterns. Theoretical guarantees include invariance, consistency, and asymptotic normality, along with a quadratic-form zero-calibration test. Experiments on circular and spherical data successfully reveal latent structures invisible to vMF-based methods, demonstrating the approach’s efficacy and practical utility.

directional uncertaintygeometric structureinformation decomposition

Hot Scholars

HZ

Hairong Zheng

Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences
biomedical imaging
PS

Pravendra Singh

Assistant Professor, IIT Roorkee
Deep LearningMachine LearningComputer VisionArtificial Intelligence
MC

Mickaël Coustaty

University of La Rochelle
Computer ScienceNLPCVDigital Humanities
JW

Jianwu Wang

Professor of Data Science, University of Maryland Baltimore County
Big Data AnalyticsEarth/Climate InformaticsCausal AIDistributed Computing
TL

Taveena Lotey

Research Scholar, Indian Institute of Technology Roorkee
Pattern recognitionDeep learningBrain-Computer Interfaces