Generalized Kernel Thinning

📅 2021-10-04
🏛️ International Conference on Learning Representations
📈 Citations: 27
✨ Influential: 2
📄 PDF
🤖 AI Summary
To address large integration errors and low posterior approximation accuracy in high-dimensional probability distribution compression, this paper proposes Generalized Kernel Thinning (GKT): a distribution compression framework operating in a reproducing kernel Hilbert space (RKHS) via flexible selection of target kernels, fractional-power kernels, or their weighted combinations. Our method establishes the first tight maximum mean discrepancy (MMD) error bound without requiring explicit square-root kernel representations; unifies treatment of both smooth and non-smooth kernels; and introduces the KT+ framework, jointly controlling function-level and distribution-level approximation errors. Experiments on 100-dimensional distributions and posterior compression for differential equation inference demonstrate significantly reduced integration error and MMD convergence rates that surpass those of Monte Carlo methods.
📝 Abstract
The kernel thinning (KT) algorithm of Dwivedi and Mackey (2021) compresses a probability distribution more effectively than independent sampling by targeting a reproducing kernel Hilbert space (RKHS) and leveraging a less smooth square-root kernel. Here we provide four improvements. First, we show that KT applied directly to the target RKHS yields tighter, dimension-free guarantees for any kernel, any distribution, and any fixed function in the RKHS. Second, we show that, for analytic kernels like Gaussian, inverse multiquadric, and sinc, target KT admits maximum mean discrepancy (MMD) guarantees comparable to or better than those of square-root KT without making explicit use of a square-root kernel. Third, we prove that KT with a fractional power kernel yields better-than-Monte-Carlo MMD guarantees for non-smooth kernels, like Laplace and Mat'ern, that do not have square-roots. Fourth, we establish that KT applied to a sum of the target and power kernels (a procedure we call KT+) simultaneously inherits the improved MMD guarantees of power KT and the tighter individual function guarantees of target KT. In our experiments with target KT and KT+, we witness significant improvements in integration error even in $100$ dimensions and when compressing challenging differential equation posteriors.
Problem

Research questions and friction points this paper is trying to address.

Probability Distribution Compression
High-Dimensional Data
Complex Kernel Functions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Kernel Trimming
Dimension-independent Guarantees
KT+ Methodology
Harvard University | MIT | Microsoft Research