mmd pseudo-posterior estimation

Design and implement inference procedures that form pseudo-posteriors by reweighting prior simulations according to kernel two-sample discrepancies (maximum mean discrepancy, MMD) between observed and simulated distributions, i.e., distributional-matching inference using MMD-based weights. Analyze and validate these MMD-based pseudo-posteriors by proving or estimating their Monte Carlo consistency and concentration properties.

mmdpseudo-posteriorestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the need for more efficient, robust, and flexible metrics for measuring distances between probability distributions in statistical inference and numerical integration. Centered on kernel methods, we propose an efficient estimator for Maximum Mean Discrepancy (MMD), develop novel MMD-based approaches for conditional expectation estimation and integral calibration, and introduce a new family of distance measures—kernel quantile discrepancies—that effectively overcome MMD’s limitations in tail sensitivity and discriminative power. Both theoretical analysis and empirical experiments demonstrate that the proposed methods offer strong scalability, computational efficiency, and superior performance, thereby providing more powerful and practical kernel-based tools for nonparametric statistics and integration tasks.

integrationkernel-based distancesmaximum mean discrepancy

Signature Maximum Mean Discrepancy Two-Sample Statistical Tests

Jun 02, 2025
AA
Andrew Alden
🏛️ King's College London | University of Oxford

This paper addresses the two-sample testing problem on path space—determining whether two sets of time-series paths originate from the same stochastic process. Method: We propose the signature Maximum Mean Discrepancy (sig-MMD), a kernel-based statistic built upon the signature transform, to quantify distributional discrepancies between path measures. Contribution/Results: We establish the first systematic theoretical framework for sig-MMD, identifying that statistical power degradation—particularly elevated Type-II error under finite samples—stems from the coupling of signature truncation and kernel bandwidth selection. To mitigate this, we introduce an adaptive truncation order selection scheme and a data-driven bandwidth correction strategy. Experiments across multiple synthetic and real-world path datasets demonstrate that our method reduces misclassification rates by 35%–62% compared to baselines, achieves superior robustness over existing kernel-based two-sample tests, and provides a reproducible, high-accuracy statistical inference tool for detecting differences among complex stochastic processes.

Addressing Type 2 errors in sig-MMD hypothesis tests with mitigation techniquesApplying sig-MMD to test if path sets originate from same stochastic processExtending MMD to compare path space distributions using signature kernel

A Permutation-free Kernel Two-Sample Test

Nov 27, 2022
SS
S. Shekhar
🏛️ Carnegie Mellon University | Yonsei University

Traditional kernel Maximum Mean Discrepancy (MMD) two-sample tests rely on permutation to determine critical thresholds, ensuring finite-sample validity but incurring an O(n²) computational cost per permutation—prohibitively expensive for large samples. This paper proposes the cross-MMD test statistic: by splitting samples to construct a U-statistic, and combining studentization with a Gaussian kernel, it yields the first kernel MMD test that requires no permutations. The method achieves asymptotic normality with a single O(n²) computation, while preserving finite-sample validity, statistical consistency, and minimax optimal detection rates under local alternatives. Theoretically and empirically, cross-MMD accelerates testing by over an order of magnitude compared to permutation-based approaches on large samples, with only a marginal loss in power, and maintains strong consistency against any fixed distributional discrepancy.

Achieves asymptotic normality and minimax optimal powerAvoids computationally expensive permutation methodsProposes cross-MMD for efficient two-sample testing

Targeted Separation and Convergence with Kernel Discrepancies

Sep 26, 2022
AB
A. Barp
🏛️ University College London | The Alan Turing Institute | Mirelo AI | University of Cambridge | Microsoft Research

Kernelized Stein discrepancy (KSD) suffers from theoretical limitations in controlling weak convergence and precisely separating target distributions. Method: We integrate Bochner embedding theory, Stein’s method, kernel analysis, and weak topology theory to systematically address these limitations. Contribution/Results: First, we establish the necessary and sufficient conditions for KSD to metrize weak convergence—its first rigorous characterization. Second, we construct a novel class of unbounded kernels that are universally discriminative—capable of separating all Borel probability measures—overcoming the inherent discriminability constraints of bounded kernels. Third, we propose the first KSD variant provably equivalent to weak convergence. Our framework significantly enhances KSD’s separation power and convergence control: on ℝᵈ, it enables precise quantitative characterization of weak convergence toward any target distribution P. This advancement strengthens theoretical guarantees and empirical performance in statistical hypothesis testing, sample quality assessment, and Stein variational gradient descent (SVGD) sampling.

Characterize kernels for separating Bochner embeddable measuresDevelop conditions for weak convergence control with bounded kernelsExpand KSD conditions to metrize convergence for hypothesis testing

Exact Sampling of Gibbs Measures with Estimated Losses

Apr 24, 2024
DF
David Frazier
🏛️ Monash University | University College London | Queensland University of Technology

This work addresses the slow MCMC convergence in Gibbs posterior sampling under stochastic loss functions, which stems from spurious dependence on the number of pseudo-observations. We propose the first pseudo-sample-size–independent corrected piecewise deterministic Markov process (PDMP) sampler. By designing a novel jump-rate function and direction mechanism, our method rigorously ensures that the invariant measure remains invariant to the pseudo-observation count—thereby overcoming the inherent trade-off between asymptotic bias and slow convergence in conventional stochastic-loss inference. We prove that the sampler converges exactly to the target Gibbs posterior measure with a uniform convergence rate independent of pseudo-sample size. Empirical validation across three canonical settings—likelihood-intractable models, misspecified models, and stochastic losses—demonstrates elimination of pseudo-sample-size bias in posterior sampling, alongside substantial improvements in robustness and estimation accuracy.

Addressing slow convergence in Gibbs measures with estimated lossesImproving inference for intractable likelihoods and model misspecificationReducing pseudo-observation dependence in MCMC posterior sampling

Latest Papers

What's happening recently
View more

This study establishes minimax lower bounds for the estimation of Maximum Mean Discrepancy (MMD), Hilbert–Schmidt Independence Criterion (HSIC), and Kernelized Stein Discrepancy (KSD) in general topological spaces under unbounded kernel conditions. By integrating reproducing kernel Hilbert space theory, functional analysis, and a minimax information-theoretic framework, the work rigorously proves—under mild assumptions—that the optimal convergence rate for these three classes of kernel-based discrepancy measures remains $n^{-1/2}$. This result resolves a long-standing open theoretical question and extends to the estimation of mean embeddings and centered cross-covariance operators, thereby establishing the minimax optimality of their parametric convergence rates.

HSICkernel discrepancyKSD

This work addresses the critical dependence of Maximum Mean Discrepancy (MMD) two-sample test power on kernel selection, a challenge exacerbated by existing data-driven approaches that either overfit due to violations of the i.i.d. assumption or fail to scale to continuous kernel spaces. The paper pioneers a rigorous formulation of kernel selection as a model selection problem and introduces the Complexity-Penalized MMD (CP-MMD) criterion. By deriving a complexity penalty from uniform concentration inequalities for two-sample statistics, CP-MMD seamlessly integrates into the optimization objective, enabling direct tuning of continuous kernel parameters—such as bandwidths, polynomial features, or even deep network weights—without requiring grid search. The method maintains strict Type I error control while achieving or surpassing state-of-the-art test power across diverse experimental settings.

Data-Driven OptimizationKernel SelectionMaximum Mean Discrepancy

This work addresses the lack of convergence guarantees for Maximum Mean Discrepancy (MMD) estimation in non-convex settings. By adopting the perspective of MMD gradient flows, the authors propose a Preconditioned Gradient Descent (PGD) algorithm that performs parameter optimization in the space of probability measures. They establish, for the first time, global asymptotic convergence of PGD under non-convexity by introducing gradient domination and projected residual conditions, thereby bridging nonparametric gradient flows with parametric optimization. Experimental results demonstrate that PGD significantly outperforms standard gradient descent in both parameter estimation and composite hypothesis testing tasks, corroborating both the theoretical rigor and practical efficacy of the proposed method.

Minimum MMD estimationnon-convexityoptimization problem

This study addresses the quantification of posterior uncertainty in kernel density estimation within a predictive Bayesian framework. By analyzing the predictive measure induced by kernel density estimators, the authors establish, for the first time, that the associated resampling sequence—despite failing to satisfy conditional independence and identical distribution (c.i.d.) or asymptotic c.i.d. (a.c.i.d.) conditions—converges weakly almost surely. In the case of Gaussian kernels, they further derive an explicit density representation of the limiting random probability measure and construct corresponding moment estimators. Leveraging these results, the paper successfully derives Bayesian credible intervals for kernel density estimates and demonstrates their empirical validity on two real-world datasets, thereby providing a rigorous tool for uncertainty quantification in nonparametric density estimation.

asymptotic exchangeabilitycredibility intervalskernel density estimation

This study investigates the construction of Bayesian predictive inference methods with favorable asymptotic properties and offers a novel Bayesian interpretation of kernel density estimation. Focusing on two classes of predictive rules—classical kernel density estimators and their recursive variants—the work systematically examines their weak almost sure convergence under sequential observations by integrating nonparametric kernel methods, stochastic process theory, and weak convergence analysis. The analysis reveals that the classical estimator converges weakly almost surely to a probability measure with compact support, whereas its recursive counterpart converges to one with non-compact support. Beyond establishing weak almost sure convergence for both schemes, this research extends the theoretical foundations of kernel methods within Bayesian predictive inference and provides a fresh Bayesian perspective on kernel density estimation.

Bayesian inferencekernel density estimationpredictive inference

Hot Scholars

SG

Sajjad Ghiasvand

PhD Student at UC Santa Barbara
Machine LearningOptimizationPEFT
QK

Qingbo Kang

West China Hospital
Deep LearningMedical Image AnalysisComputer Vision
MA

Mahnoosh Alizadeh

Associate Professor of Electrical and Computer Engineering, University of California Santa Barbara
NetworksCyber-Physical SystemsSmart GridLearning
ZW

Zhipeng Wei

ICSI, UC Berkeley
robustness of deep learning