Score
Design and implement inference procedures that form pseudo-posteriors by reweighting prior simulations according to kernel two-sample discrepancies (maximum mean discrepancy, MMD) between observed and simulated distributions, i.e., distributional-matching inference using MMD-based weights. Analyze and validate these MMD-based pseudo-posteriors by proving or estimating their Monte Carlo consistency and concentration properties.
This work addresses the need for more efficient, robust, and flexible metrics for measuring distances between probability distributions in statistical inference and numerical integration. Centered on kernel methods, we propose an efficient estimator for Maximum Mean Discrepancy (MMD), develop novel MMD-based approaches for conditional expectation estimation and integral calibration, and introduce a new family of distance measures—kernel quantile discrepancies—that effectively overcome MMD’s limitations in tail sensitivity and discriminative power. Both theoretical analysis and empirical experiments demonstrate that the proposed methods offer strong scalability, computational efficiency, and superior performance, thereby providing more powerful and practical kernel-based tools for nonparametric statistics and integration tasks.
This paper addresses the two-sample testing problem on path space—determining whether two sets of time-series paths originate from the same stochastic process. Method: We propose the signature Maximum Mean Discrepancy (sig-MMD), a kernel-based statistic built upon the signature transform, to quantify distributional discrepancies between path measures. Contribution/Results: We establish the first systematic theoretical framework for sig-MMD, identifying that statistical power degradation—particularly elevated Type-II error under finite samples—stems from the coupling of signature truncation and kernel bandwidth selection. To mitigate this, we introduce an adaptive truncation order selection scheme and a data-driven bandwidth correction strategy. Experiments across multiple synthetic and real-world path datasets demonstrate that our method reduces misclassification rates by 35%–62% compared to baselines, achieves superior robustness over existing kernel-based two-sample tests, and provides a reproducible, high-accuracy statistical inference tool for detecting differences among complex stochastic processes.
Traditional kernel Maximum Mean Discrepancy (MMD) two-sample tests rely on permutation to determine critical thresholds, ensuring finite-sample validity but incurring an O(n²) computational cost per permutation—prohibitively expensive for large samples. This paper proposes the cross-MMD test statistic: by splitting samples to construct a U-statistic, and combining studentization with a Gaussian kernel, it yields the first kernel MMD test that requires no permutations. The method achieves asymptotic normality with a single O(n²) computation, while preserving finite-sample validity, statistical consistency, and minimax optimal detection rates under local alternatives. Theoretically and empirically, cross-MMD accelerates testing by over an order of magnitude compared to permutation-based approaches on large samples, with only a marginal loss in power, and maintains strong consistency against any fixed distributional discrepancy.
Kernelized Stein discrepancy (KSD) suffers from theoretical limitations in controlling weak convergence and precisely separating target distributions. Method: We integrate Bochner embedding theory, Stein’s method, kernel analysis, and weak topology theory to systematically address these limitations. Contribution/Results: First, we establish the necessary and sufficient conditions for KSD to metrize weak convergence—its first rigorous characterization. Second, we construct a novel class of unbounded kernels that are universally discriminative—capable of separating all Borel probability measures—overcoming the inherent discriminability constraints of bounded kernels. Third, we propose the first KSD variant provably equivalent to weak convergence. Our framework significantly enhances KSD’s separation power and convergence control: on ℝᵈ, it enables precise quantitative characterization of weak convergence toward any target distribution P. This advancement strengthens theoretical guarantees and empirical performance in statistical hypothesis testing, sample quality assessment, and Stein variational gradient descent (SVGD) sampling.
This work addresses the slow MCMC convergence in Gibbs posterior sampling under stochastic loss functions, which stems from spurious dependence on the number of pseudo-observations. We propose the first pseudo-sample-size–independent corrected piecewise deterministic Markov process (PDMP) sampler. By designing a novel jump-rate function and direction mechanism, our method rigorously ensures that the invariant measure remains invariant to the pseudo-observation count—thereby overcoming the inherent trade-off between asymptotic bias and slow convergence in conventional stochastic-loss inference. We prove that the sampler converges exactly to the target Gibbs posterior measure with a uniform convergence rate independent of pseudo-sample size. Empirical validation across three canonical settings—likelihood-intractable models, misspecified models, and stochastic losses—demonstrates elimination of pseudo-sample-size bias in posterior sampling, alongside substantial improvements in robustness and estimation accuracy.
This study establishes minimax lower bounds for the estimation of Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernelized Stein Discrepancy (KSD) in general topological spaces under unbounded kernel conditions. By integrating reproducing kernel Hilbert space theory, functional analysis, and a minimax information-theoretic framework, the work rigorously proves—under mild assumptions—that the optimal convergence rate for these three classes of kernel-based discrepancy measures remains $n^{-1/2}$. This result resolves a long-standing open theoretical question and extends to the estimation of mean embeddings and centered cross-covariance operators, thereby establishing the minimax optimality of their parametric convergence rates.
This work addresses the critical dependence of Maximum Mean Discrepancy (MMD) two-sample test power on kernel selection, a challenge exacerbated by existing data-driven approaches that either overfit due to violations of the i.i.d. assumption or fail to scale to continuous kernel spaces. The paper pioneers a rigorous formulation of kernel selection as a model selection problem and introduces the Complexity-Penalized MMD (CP-MMD) criterion. By deriving a complexity penalty from uniform concentration inequalities for two-sample statistics, CP-MMD seamlessly integrates into the optimization objective, enabling direct tuning of continuous kernel parameters—such as bandwidths, polynomial features, or even deep network weights—without requiring grid search. The method maintains strict Type I error control while achieving or surpassing state-of-the-art test power across diverse experimental settings.
This work addresses the lack of convergence guarantees for Maximum Mean Discrepancy (MMD) estimation in non-convex settings. By adopting the perspective of MMD gradient flows, the authors propose a Preconditioned Gradient Descent (PGD) algorithm that performs parameter optimization in the space of probability measures. They establish, for the first time, global asymptotic convergence of PGD under non-convexity by introducing gradient domination and projected residual conditions, thereby bridging nonparametric gradient flows with parametric optimization. Experimental results demonstrate that PGD significantly outperforms standard gradient descent in both parameter estimation and composite hypothesis testing tasks, corroborating both the theoretical rigor and practical efficacy of the proposed method.
This study addresses the quantification of posterior uncertainty in kernel density estimation within a predictive Bayesian framework. By analyzing the predictive measure induced by kernel density estimators, the authors establish, for the first time, that the associated resampling sequence—despite failing to satisfy conditional independence and identical distribution (c.i.d.) or asymptotic c.i.d. (a.c.i.d.) conditions—converges weakly almost surely. In the case of Gaussian kernels, they further derive an explicit density representation of the limiting random probability measure and construct corresponding moment estimators. Leveraging these results, the paper successfully derives Bayesian credible intervals for kernel density estimates and demonstrates their empirical validity on two real-world datasets, thereby providing a rigorous tool for uncertainty quantification in nonparametric density estimation.
This study investigates the construction of Bayesian predictive inference methods with favorable asymptotic properties and offers a novel Bayesian interpretation of kernel density estimation. Focusing on two classes of predictive rules—classical kernel density estimators and their recursive variants—the work systematically examines their weak almost sure convergence under sequential observations by integrating nonparametric kernel methods, stochastic process theory, and weak convergence analysis. The analysis reveals that the classical estimator converges weakly almost surely to a probability measure with compact support, whereas its recursive counterpart converges to one with non-compact support. Beyond establishing weak almost sure convergence for both schemes, this research extends the theoretical foundations of kernel methods within Bayesian predictive inference and provides a fresh Bayesian perspective on kernel density estimation.