random-feature mmd approximation

Designs and implements approximations of the Maximum Mean Discrepancy (MMD) statistic using finite random feature maps (such as random Fourier features), producing low-dimensional feature representations that approximate kernel evaluations. Builds algorithms and error analyses that reduce MMD computation to linear time and enable efficient evaluation and gradient computation through the feature-based approximation.

random-featuremmdapproximation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Signature Maximum Mean Discrepancy Two-Sample Statistical Tests

Jun 02, 2025
AA
Andrew Alden
🏛️ King's College London | University of Oxford

This paper addresses the two-sample testing problem on path space—determining whether two sets of time-series paths originate from the same stochastic process. Method: We propose the signature Maximum Mean Discrepancy (sig-MMD), a kernel-based statistic built upon the signature transform, to quantify distributional discrepancies between path measures. Contribution/Results: We establish the first systematic theoretical framework for sig-MMD, identifying that statistical power degradation—particularly elevated Type-II error under finite samples—stems from the coupling of signature truncation and kernel bandwidth selection. To mitigate this, we introduce an adaptive truncation order selection scheme and a data-driven bandwidth correction strategy. Experiments across multiple synthetic and real-world path datasets demonstrate that our method reduces misclassification rates by 35%–62% compared to baselines, achieves superior robustness over existing kernel-based two-sample tests, and provides a reproducible, high-accuracy statistical inference tool for detecting differences among complex stochastic processes.

Addressing Type 2 errors in sig-MMD hypothesis tests with mitigation techniquesApplying sig-MMD to test if path sets originate from same stochastic processExtending MMD to compare path space distributions using signature kernel

This work addresses the inconsistent performance of existing variance estimators for the Maximum Mean Discrepancy (MMD) two-sample test under varying conditions—specifically, across the null and alternative hypotheses as well as balanced and unbalanced sample settings—and the absence of a unified framework. By leveraging the U-statistic representation and Hoeffding decomposition, the authors establish the first unified, unbiased variance estimation framework for MMD that encompasses all such hypothesis and sampling configurations. Furthermore, for the one-dimensional Laplacian kernel, they develop an exact accelerated algorithm that reduces computational complexity from O(n²) to O(n log n). The proposed method demonstrates robustness in finite samples, significantly enhancing both statistical inference accuracy and computational efficiency.

Imbalanced DataMaximum Mean DiscrepancyTwo-sample Testing

A Practical Introduction to Kernel Discrepancies: MMD, HSIC&KSD

Mar 04, 2025
AS
Antonin Schrab
🏛️ University College London

This work addresses the kernel selection and bandwidth sensitivity challenges inherent in kernel-based discrepancy measures—specifically Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernel Stein Discrepancy (KSD)—for distribution comparison, independence testing, and generative model evaluation. We propose a unified computational framework and a multi-kernel adaptive fusion estimator grounded in Hilbert space embeddings and Stein operator theory. Our method integrates V- and U-statistics, employs efficient incomplete U-statistic approximations, and incorporates a data-driven bandwidth adaptation strategy. Compared to single-kernel approaches, the proposed estimator substantially improves statistical power in small-sample and high-dimensional settings, while ensuring reproducibility and ease of hyperparameter tuning. The resulting toolkit provides a theoretically coherent and practically accessible unified implementation for all three major kernel discrepancies.

Addresses kernel selection and bandwidth impact.Introduces kernel discrepancies: MMD, HSIC, KSD.Presents estimators like V-statistics, U-statistics.

Finite sample properties of parametric MMD estimation: robustness to misspecification and dependence

Dec 12, 2019
BC
Badr-Eddine Chérief-Abdellatif
🏛️ CNRS | ESSEC Business School

This paper addresses the general parametric estimation problem without distributional assumptions, aiming to construct estimators robust to both model misspecification and complex data dependence structures. We propose a minimum distance estimation framework based on the Maximum Mean Discrepancy (MMD), establishing— for the first time—its statistical consistency under non-i.i.d. sampling and model misspecification, and providing theoretical convergence guarantees for stochastic gradient descent optimization. Theoretically, the estimator exhibits intrinsic robustness to temporal dependence, outliers, and distributional shifts. Numerical experiments demonstrate its superior performance over classical M-estimators under data contamination and intricate time-series settings. Our core contribution lies in unifying the characterization of the MMD estimator’s generalization error and robustness limits, thereby offering a novel paradigm for reliable inference with nonstandard data.

Assesses estimator robustness to dependence and outliersDevelops a universal estimation procedure using MMDTheoretical study of SGD algorithm for estimator computation

A Permutation-free Kernel Two-Sample Test

Nov 27, 2022
SS
S. Shekhar
🏛️ Carnegie Mellon University | Yonsei University

Traditional kernel Maximum Mean Discrepancy (MMD) two-sample tests rely on permutation to determine critical thresholds, ensuring finite-sample validity but incurring an O(n²) computational cost per permutation—prohibitively expensive for large samples. This paper proposes the cross-MMD test statistic: by splitting samples to construct a U-statistic, and combining studentization with a Gaussian kernel, it yields the first kernel MMD test that requires no permutations. The method achieves asymptotic normality with a single O(n²) computation, while preserving finite-sample validity, statistical consistency, and minimax optimal detection rates under local alternatives. Theoretically and empirically, cross-MMD accelerates testing by over an order of magnitude compared to permutation-based approaches on large samples, with only a marginal loss in power, and maintains strong consistency against any fixed distributional discrepancy.

Achieves asymptotic normality and minimax optimal powerAvoids computationally expensive permutation methodsProposes cross-MMD for efficient two-sample testing

Latest Papers

What's happening recently
View more

This study establishes minimax lower bounds for the estimation of Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernelized Stein Discrepancy (KSD) in general topological spaces under unbounded kernel conditions. By integrating reproducing kernel Hilbert space theory, functional analysis, and a minimax information-theoretic framework, the work rigorously proves—under mild assumptions—that the optimal convergence rate for these three classes of kernel-based discrepancy measures remains $n^{-1/2}$. This result resolves a long-standing open theoretical question and extends to the estimation of mean embeddings and centered cross-covariance operators, thereby establishing the minimax optimality of their parametric convergence rates.

HSICkernel discrepancyKSD

This work addresses the critical dependence of Maximum Mean Discrepancy (MMD) two-sample test power on kernel selection, a challenge exacerbated by existing data-driven approaches that either overfit due to violations of the i.i.d. assumption or fail to scale to continuous kernel spaces. The paper pioneers a rigorous formulation of kernel selection as a model selection problem and introduces the Complexity-Penalized MMD (CP-MMD) criterion. By deriving a complexity penalty from uniform concentration inequalities for two-sample statistics, CP-MMD seamlessly integrates into the optimization objective, enabling direct tuning of continuous kernel parameters—such as bandwidths, polynomial features, or even deep network weights—without requiring grid search. The method maintains strict Type I error control while achieving or surpassing state-of-the-art test power across diverse experimental settings.

Data-Driven OptimizationKernel SelectionMaximum Mean Discrepancy

This work addresses the lack of convergence guarantees for Maximum Mean Discrepancy (MMD) estimation in non-convex settings. By adopting the perspective of MMD gradient flows, the authors propose a Preconditioned Gradient Descent (PGD) algorithm that performs parameter optimization in the space of probability measures. They establish, for the first time, global asymptotic convergence of PGD under non-convexity by introducing gradient domination and projected residual conditions, thereby bridging nonparametric gradient flows with parametric optimization. Experimental results demonstrate that PGD significantly outperforms standard gradient descent in both parameter estimation and composite hypothesis testing tasks, corroborating both the theoretical rigor and practical efficacy of the proposed method.

Minimum MMD estimationnon-convexityoptimization problem

This work addresses the computational and memory bottlenecks of traditional Grassmannian kernel methods, which require constructing full Gram matrices and thus struggle with high-dimensional subspace data. To overcome these limitations, the authors propose a scalable kernel approximation framework based on random rank-one projections combined with bounded nonlinear transformations—either periodic or binary—that yield compact one-bit subspace feature representations. This approach enables continuous interpolation between the inverse Binet–Cauchy kernel and Gaussian-like kernels while effectively preserving the intrinsic geometry of subspaces. The method substantially reduces computational, memory, and storage costs. Experimental results on synthetic data and the ETH-80 classification benchmark demonstrate that the proposed technique accurately maintains Grassmannian geometric relationships with high fidelity, confirming its efficiency and practical utility.

Grassmannian kernelsrandom featuresrank-one projections

Hot Scholars

SF

Soukaina Filali Boubrahimi

Associate Professor of Computer Science, Utah State University, Logan, Utah, USA
Time Series Data MiningDatabase SystemsData MiningSpatio-Temporal Pattern Mining
YX

Yinhao Xiao

Guangdong University of Finance and Economics
System Security
TD

Tommaso Dorigo

First Researcher, INFN - Sezione di Padova
Particle physicsastrophysicsmachine learningstatistical data analysis
ME

MohammadReza EskandariNasab

PhD Candidate in the School of Computing, Utah State University.
Generative ModelsTime Series AnalysisSolar Flare PredictionBiomedical Signal Processing
NB

Nadjet Bourdache

GREYC, Université de Caen Normandie
Algorithmic Decision Theoryartificial intelligenceoperational research