From Continuous Dynamics to Practical Gradient-Based Samplers

📅 2026-08-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiency and lack of a unified theoretical understanding of gradient-based MCMC methods when sampling from anisotropic and hierarchical posteriors. Starting from continuous-time dynamics, the authors systematically derive algorithms such as HMC, MALA, NUTS, and MAKLA through numerical discretization and Metropolis correction, establishing a cohesive theoretical framework. They introduce two key innovations: a globally whitened mass matrix and a stochastic step-size strategy, which together mitigate sampling difficulties arising from state-dependent curvature. The proposed approach substantially improves sampling efficiency and demonstrates superior convergence and stability in large-scale Bayesian inference and hierarchical models.
📝 Abstract
Gradient-based Markov chain Monte Carlo methods are often introduced as a catalog of algorithms: Hamiltonian Monte Carlo (HMC), the Metropolis-adjusted Langevin algorithm (MALA), the No-U-Turn Sampler (NUTS), and several underdamped variants. This presentation obscures the common structure of the methods and, more importantly, the reasons why a sampler that is correct in principle may be ineffective in practice. We develop a unified account, beginning with exact continuous-time dynamics that represent idealized sampling methods and for which Metropolis adjustments are not required. Numerical discretization makes the dynamics computationally feasible but introduces bias. Metropolis adjustment removes the asymptotic bias by converting numerical errors into rejection, leading to HMC, MALA, NUTS, and the Metropolis-adjusted kinetic Langevin algorithm (MAKLA). The second half of the paper presents geometric design choices that determine practical performance, namely, although MAKLA and NUTS have nice theoretical properties, their sampling efficiency may be slow in practice. Importantly, a fixed mass matrix can whiten globally anisotropic targets, often fixing sampling inefficiency in Bayesian posteriors with large data. Whereas hierarchical posteriors introduce their own problem, causing state-dependent variation in the Hessian (e.g., Neal's funnel). We explain how a randomized step size can be used effectively to sample from such a distribution. The resulting paper is both a tutorial on the mechanics of gradient-based sampling and a set of practical recipes to improve sampler performance.
Problem

Research questions and friction points this paper is trying to address.

gradient-based MCMC
sampling efficiency
anisotropic targets
hierarchical posteriors
state-dependent Hessian
Innovation

Methods, ideas, or system contributions that make the work stand out.

gradient-based MCMC
continuous-time dynamics
Metropolis adjustment
mass matrix whitening
randomized step size
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
James Chok
School of Mathematics and Maxwell Institute for Mathematical Sciences, The University of Edinburgh, Edinburgh, EH9 3FD, United Kingdom