Score
Design and implement methods that detect, quantify, and localize differences between probability distributions or embedding spaces using kernel-based discrepancy measures such as kernel mean embeddings and Maximum Mean Discrepancy (MMD). This includes selecting and tuning kernels and bandwidths, estimating kernel mean embeddings and MMD test statistics, identifying interpretable discrepancy directions and subsets, and applying contrastive embedding clustering or discrepancy modeling to discover or suppress coherent structures across datasets.
This study addresses the critical need for quantifying dataset similarity in model generalization, transfer learning, simulation calibration, and two-sample testing. We systematically survey 118 similarity quantification methods and propose the first ten-dimensional classification framework, organizing approaches into seven technical categories: statistical distances (e.g., Wasserstein distance, Maximum Mean Discrepancy), kernel-based methods, information-theoretic measures, dimensionality-reduction embeddings, permutation tests, generative-model-based discriminators, and Gaussian process likelihood ratios. We develop a multi-dimensional evaluation system balancing theoretical guarantees, interpretability, and practical applicability, yielding a structured recommendation matrix aligned with task requirements and data characteristics. Furthermore, we introduce the first open-source, interactive tool for method selection—enabling real-time filtering and parameter configuration—to significantly enhance both selection efficiency and deployment suitability.
This work addresses the need for more efficient, robust, and flexible metrics for measuring distances between probability distributions in statistical inference and numerical integration. Centered on kernel methods, we propose an efficient estimator for Maximum Mean Discrepancy (MMD), develop novel MMD-based approaches for conditional expectation estimation and integral calibration, and introduce a new family of distance measures—kernel quantile discrepancies—that effectively overcome MMD’s limitations in tail sensitivity and discriminative power. Both theoretical analysis and empirical experiments demonstrate that the proposed methods offer strong scalability, computational efficiency, and superior performance, thereby providing more powerful and practical kernel-based tools for nonparametric statistics and integration tasks.
This work addresses the kernel selection and bandwidth sensitivity challenges inherent in kernel-based discrepancy measures—specifically Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernel Stein Discrepancy (KSD)—for distribution comparison, independence testing, and generative model evaluation. We propose a unified computational framework and a multi-kernel adaptive fusion estimator grounded in Hilbert space embeddings and Stein operator theory. Our method integrates V- and U-statistics, employs efficient incomplete U-statistic approximations, and incorporates a data-driven bandwidth adaptation strategy. Compared to single-kernel approaches, the proposed estimator substantially improves statistical power in small-sample and high-dimensional settings, while ensuring reproducibility and ease of hyperparameter tuning. The resulting toolkit provides a theoretically coherent and practically accessible unified implementation for all three major kernel discrepancies.
Kernelized Stein discrepancy (KSD) suffers from theoretical limitations in controlling weak convergence and precisely separating target distributions. Method: We integrate Bochner embedding theory, Stein’s method, kernel analysis, and weak topology theory to systematically address these limitations. Contribution/Results: First, we establish the necessary and sufficient conditions for KSD to metrize weak convergence—its first rigorous characterization. Second, we construct a novel class of unbounded kernels that are universally discriminative—capable of separating all Borel probability measures—overcoming the inherent discriminability constraints of bounded kernels. Third, we propose the first KSD variant provably equivalent to weak convergence. Our framework significantly enhances KSD’s separation power and convergence control: on ℝᵈ, it enables precise quantitative characterization of weak convergence toward any target distribution P. This advancement strengthens theoretical guarantees and empirical performance in statistical hypothesis testing, sample quality assessment, and Stein variational gradient descent (SVGD) sampling.
To address the $O(n^2)$ computational bottleneck of kernelized Stein discrepancy (KSD) under large-scale data—arising from its reliance on U- or V-statistics—this paper introduces, for the first time, the Nyström low-rank kernel approximation into KSD estimation, yielding a scalable and accelerated KSD estimator. The proposed method reduces time complexity to $O(mn + m^3)$, where $m ll n$, and establishes $sqrt{n}$-consistency under sub-Gaussian assumptions. Theoretical analysis is grounded in the Stein operator and reproducing kernel Hilbert space (RKHS) framework, balancing statistical efficiency with computational tractability. Extensive benchmark experiments demonstrate that the new estimator retains statistical power comparable to the original KSD while substantially enhancing practicality for large-scale goodness-of-fit testing. This work provides an efficient, theoretically sound tool for high-dimensional distribution fitting and hypothesis testing.
Traditional kernel Maximum Mean Discrepancy (MMD) two-sample tests rely on permutation to determine critical thresholds, ensuring finite-sample validity but incurring an O(n²) computational cost per permutation—prohibitively expensive for large samples. This paper proposes the cross-MMD test statistic: by splitting samples to construct a U-statistic, and combining studentization with a Gaussian kernel, it yields the first kernel MMD test that requires no permutations. The method achieves asymptotic normality with a single O(n²) computation, while preserving finite-sample validity, statistical consistency, and minimax optimal detection rates under local alternatives. Theoretically and empirically, cross-MMD accelerates testing by over an order of magnitude compared to permutation-based approaches on large samples, with only a marginal loss in power, and maintains strong consistency against any fixed distributional discrepancy.
This work addresses the fundamental challenge of measuring discrepancies between conditional distributions in statistics and machine learning. The authors propose Conditional Maximum Mean Discrepancy (CMMD), a unified framework grounded in reproducing kernel Hilbert space embeddings, which establishes a hierarchical family ranging from CMMD₀ to CMMDₛ and elucidates their intrinsic mathematical relationships. A key innovation is the introduction of a doubly robust estimator that guarantees consistent estimation as long as either the conditional mean model or the weighting model is correctly specified. Combining operator smoothing with nonparametric techniques, the theoretical analysis is complemented by empirical results demonstrating that CMMD effectively captures complex conditional dependence structures and significantly outperforms existing methods in conditional distribution testing tasks.
This study establishes minimax lower bounds for the estimation of Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernelized Stein Discrepancy (KSD) in general topological spaces under unbounded kernel conditions. By integrating reproducing kernel Hilbert space theory, functional analysis, and a minimax information-theoretic framework, the work rigorously proves—under mild assumptions—that the optimal convergence rate for these three classes of kernel-based discrepancy measures remains $n^{-1/2}$. This result resolves a long-standing open theoretical question and extends to the estimation of mean embeddings and centered cross-covariance operators, thereby establishing the minimax optimality of their parametric convergence rates.
This study addresses the challenge of statistical inference in the presence of heteroscedastic measurement errors, where conventional methods often fail due to either neglecting the noise structure or incurring prohibitive computational costs. The authors propose a convolutional Maximum Mean Discrepancy (convMMD) framework, which extends MMD to noisy settings for the first time by convolving observed samples with the known noise distribution, thereby enabling robust nonparametric testing and estimation. Theoretical contributions include finite-sample bias bounds, an equivalence between noise-convolved MMD tests and kernel smoothing, and proofs of consistency and asymptotic normality of the resulting estimators. Empirical evaluations demonstrate that the method achieves both computational efficiency and practical utility across simulations and real-world applications in astronomy and social sciences.
This study addresses the NP-hard problem of selecting low-discrepancy subsets from large-scale sets, which has significant applications in quasi-Monte Carlo methods, machine learning, and computer graphics. We establish for the first time the NP-hardness of this problem under kernel discrepancy measures and propose a novel framework based on Bayesian optimization. By constructing a surrogate model using deep embedded kernels, our approach efficiently searches for optimal subsets, overcoming the computational bottlenecks inherent in traditional combinatorial optimization. Extensive experiments demonstrate that the method substantially reduces subset discrepancy across multiple discrepancy metrics, highlighting its effectiveness and versatility in low-discrepancy design tasks.
This work addresses the limited power of nonparametric two-sample tests in high-dimensional or complex distributional settings by proposing the spectrally truncated normalized Maximum Mean Discrepancy (st-nMMD). Built upon embeddings in a reproducing kernel Hilbert space, st-nMMD integrates covariance operator normalization with spectral truncation regularization to substantially enhance test power. The paper establishes, for the first time, a non-asymptotic exponential upper bound for st-nMMD under the null hypothesis, introduces an adaptive hyperparameter tuning algorithm that avoids data splitting, and provides explicit non-asymptotic quantile estimates. Empirical results demonstrate that the method maintains proper Type I error control while achieving superior statistical power and stability under the alternative hypothesis, significantly outperforming existing kernel-based two-sample tests.