kernel stein discrepancy estimation

Design and analyze estimators that compute the kernel Stein discrepancy (KSD) between probability measures using reproducing-kernel operators, including procedures to estimate KSD from finite samples and proofs of their error behavior. Establish and compare minimax upper and lower bounds on estimation risk, relate risk to operator quantities such as spectral risk constants and Hilbert–Schmidt (HS) norms, and analyze scaling differences (trace versus HS) that determine optimality.

kernelsteindiscrepancyestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Nyström Kernel Stein Discrepancy

Jun 12, 2024
FK
Florian Kalinke
🏛️ Karlsruhe Institute of Technology | London School of Economics | The Pennsylvania State University

To address the $O(n^2)$ computational bottleneck of kernelized Stein discrepancy (KSD) under large-scale data—arising from its reliance on U- or V-statistics—this paper introduces, for the first time, the Nyström low-rank kernel approximation into KSD estimation, yielding a scalable and accelerated KSD estimator. The proposed method reduces time complexity to $O(mn + m^3)$, where $m ll n$, and establishes $sqrt{n}$-consistency under sub-Gaussian assumptions. Theoretical analysis is grounded in the Stein operator and reproducing kernel Hilbert space (RKHS) framework, balancing statistical efficiency with computational tractability. Extensive benchmark experiments demonstrate that the new estimator retains statistical power comparable to the original KSD while substantially enhancing practicality for large-scale goodness-of-fit testing. This work provides an efficient, theoretically sound tool for high-dimensional distribution fitting and hypothesis testing.

Accelerates Kernel Stein Discrepancy computationEnsures consistency in goodness-of-fit testingReduces runtime complexity for large-scale data

Targeted Separation and Convergence with Kernel Discrepancies

Sep 26, 2022
AB
A. Barp
🏛️ University College London | The Alan Turing Institute | Mirelo AI | University of Cambridge | Microsoft Research

Kernelized Stein discrepancy (KSD) suffers from theoretical limitations in controlling weak convergence and precisely separating target distributions. Method: We integrate Bochner embedding theory, Stein’s method, kernel analysis, and weak topology theory to systematically address these limitations. Contribution/Results: First, we establish the necessary and sufficient conditions for KSD to metrize weak convergence—its first rigorous characterization. Second, we construct a novel class of unbounded kernels that are universally discriminative—capable of separating all Borel probability measures—overcoming the inherent discriminability constraints of bounded kernels. Third, we propose the first KSD variant provably equivalent to weak convergence. Our framework significantly enhances KSD’s separation power and convergence control: on ℝᵈ, it enables precise quantitative characterization of weak convergence toward any target distribution P. This advancement strengthens theoretical guarantees and empirical performance in statistical hypothesis testing, sample quality assessment, and Stein variational gradient descent (SVGD) sampling.

Characterize kernels for separating Bochner embeddable measuresDevelop conditions for weak convergence control with bounded kernelsExpand KSD conditions to metrize convergence for hypothesis testing

Controlling Moments with Kernel Stein Discrepancies

Nov 10, 2022
HK
Heishiro Kanagawa
🏛️ Newcastle University | UCL | Microsoft Research

Kernel Stein Discrepancy (KSD) lacks theoretical guarantees for moment convergence and is only valid under weak convergence, limiting its applicability in distribution approximation tasks requiring higher-order moment control. Method: We propose Diffusion KSD—a novel discrepancy measure integrating the Stein operator, reproducing kernel Hilbert space (RKHS), and diffusion process theory—to rigorously characterize *q*-Wasserstein convergence. Contribution/Results: Diffusion KSD is the first KSD variant that simultaneously ensures weak convergence and convergence of all moments up to order *q*. We establish sufficient conditions under which KSD controls moment convergence, thereby extending the theoretical convergence boundary of classical KSD. The method provides a computationally tractable yet theoretically rigorous tool for MCMC diagnostic analysis and fitting unnormalized models. Its principled foundation in optimal transport and Stein methodology enables reliable assessment of distributional approximation quality, with broad implications for Bayesian inference, generative modeling, and variational inference.

Kernel Stein DiscrepanciesMCMC sampling accuracystatistical model adequacy

On the Pinsker bound of inner product kernel regression in large dimensions

Sep 02, 2024
WL
Weihao Lu
🏛️ Tsinghua University | Peking University

This paper establishes the Pinsker lower bound for inner-product kernel regression on the high-dimensional sphere $mathbb{S}^d$, under the precise asymptotic regime where the sample size satisfies $n = alpha d^gamma (1 + o_d(1))$. Methodologically, it integrates tools from high-dimensional probability, spherical harmonic analysis, kernel method theory, minimax statistical inference, spectral theory, and asymptotics of orthogonal polynomials. The main contribution is the first derivation—within this high-dimensional asymptotic framework—of the **exact Pinsker constant** and a **minimax risk expression with explicit constants**. Unlike prior work that only characterized convergence rates, this result precisely quantifies how the optimal estimation risk depends on the dimension $d$, the sample scaling parameters $alpha$ and $gamma$, and the structural properties of the kernel function. The findings thus provide a rigorous theoretical benchmark and quantitative guidance for nonparametric inference in high dimensions.

Analyzing minimax risk in large-dimensional settingsDetermining Pinsker bound for inner product kernel regressionIdentifying exact Pinsker constant for excess risk

Statistical and Geometrical properties of regularized Kernel Kullback-Leibler divergence

Aug 29, 2024
CC
Clémentine Chazal
🏛️ CREST | ENSAE | IP Paris | INRIA | Ecole Normale Supérieure | PSL Research University

The original kernelized Kullback–Leibler (KL) divergence is ill-defined when the supports of compared distributions are disjoint—a fundamental limitation. To address this, the paper proposes a Tikhonov-regularized kernel KL divergence, constructed via covariance operator embeddings in a reproducing kernel Hilbert space (RKHS). This metric is well-defined for arbitrary probability distributions—including discrete, continuous, and mutually singular ones—and provides theoretical guarantees: a bias bound relative to the true KL divergence, finite-sample convergence rates, and a closed-form solution for discrete distributions. Furthermore, the authors formulate a Wasserstein gradient flow optimization framework for the proposed divergence, ensuring theoretical convergence, and design an efficient algorithm applicable to discrete structures such as point clouds. Experiments on point cloud transport tasks demonstrate that the method outperforms existing kernelized and Wasserstein-based approaches, achieving superior stability and robustness.

Defines regularized KKL divergence for all distributionsDerives Wasserstein gradient descent for discrete distributionsProvides bounds for regularized KKL divergence deviation

Latest Papers

What's happening recently
View more

This study addresses the minimax optimal estimation of Kernel Stein Discrepancy (KSD) when only the score function of the target distribution and a finite sample are available. By analyzing the spectral properties of the Stein covariance operator, the authors establish for the first time that the minimax risk of KSD estimation is governed by the Hilbert–Schmidt norm of this operator, yielding an optimal convergence rate of √(|C⋆|_HS / n). They propose a positive-part square-root U-statistic estimator that achieves this optimal rate. In contrast, the conventional V-statistic estimator attains only √(tr(C⋆) / n), resulting in an exponentially larger error gap in high-dimensional Gaussian settings, thereby highlighting the substantial advantage of the proposed method.

Hilbert-Schmidt normKernel Stein Discrepancyminimax estimation

This study establishes minimax lower bounds for the estimation of Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernelized Stein Discrepancy (KSD) in general topological spaces under unbounded kernel conditions. By integrating reproducing kernel Hilbert space theory, functional analysis, and a minimax information-theoretic framework, the work rigorously proves—under mild assumptions—that the optimal convergence rate for these three classes of kernel-based discrepancy measures remains $n^{-1/2}$. This result resolves a long-standing open theoretical question and extends to the estimation of mean embeddings and centered cross-covariance operators, thereby establishing the minimax optimality of their parametric convergence rates.

HSICkernel discrepancyKSD

This work addresses the limited power of nonparametric two-sample tests in high-dimensional or complex distributional settings by proposing the spectrally truncated normalized Maximum Mean Discrepancy (st-nMMD). Built upon embeddings in a reproducing kernel Hilbert space, st-nMMD integrates covariance operator normalization with spectral truncation regularization to substantially enhance test power. The paper establishes, for the first time, a non-asymptotic exponential upper bound for st-nMMD under the null hypothesis, introduces an adaptive hyperparameter tuning algorithm that avoids data splitting, and provides explicit non-asymptotic quantile estimates. Empirical results demonstrate that the method maintains proper Type I error control while achieving superior statistical power and stability under the alternative hypothesis, significantly outperforming existing kernel-based two-sample tests.

kernel methodsnon-asymptotic analysisnormalized MMD

This work addresses the computational inefficiency of traditional kernel Stein discrepancy (KSD) tests, which suffer from quadratic time complexity due to their reliance on U- or V-statistics and require computationally intensive bootstrap procedures to approximate the null distribution. To overcome these limitations, the authors propose an accelerated KSD test based on the Nyström approximation. They provide the first theoretical guarantee that this approach preserves asymptotic type-I error control and local consistency within a bootstrap framework, while substantially reducing computational cost. Empirical evaluations on spherical and functional data demonstrate that the accelerated method achieves statistical performance comparable to the original KSD test but with significantly improved computational efficiency, thereby enabling scalable and theoretically sound nonparametric goodness-of-fit testing.

BootstrapComputational EfficiencyGoodness-of-Fit Test

This work addresses the exponential decay of signal-to-noise ratio (SNR) in classical polynomial Stein discrepancies when increasing the polynomial order, which severely undermines the statistical power of goodness-of-fit tests due to uncontrolled variance. For the first time, the construction of Stein discrepancies is explicitly formulated as an SNR² maximization problem. The authors propose the λ-PSD method, which integrates covariance-aware reweighting, low-dimensional subspace approximation, and Rayleigh quotient–based optimization of Stein eigenfunctions. Under Gaussian assumptions, this approach effectively prevents SNR collapse induced by high-order polynomials while retaining linear time complexity. Empirical results demonstrate substantially improved test power, highlighting the critical role of SNR-aware design in scalable Stein discrepancy methods.

Goodness-of-Fit TestingPolynomial Stein DiscrepancyScalable Inference

Hot Scholars

SB

Sayan Banerjee

Associate Professor of Statistics, University of North Carolina, Chapel Hill
Probability Theory
DZ

Dongmian Zou

Duke Kunshan University
applied harmonic analysismachine learning
HS

Hassan Sajjad

Faculty of Computer Science, Dalhousie University
Deep LearningNLPInterpretabilityExplainableAI
AC

Ashia C. Wilson

Assistant Professor at MIT
Machine LearningOptimizationDynamical SystemsStatistics