🤖 AI Summary
This work addresses the exponential convergence of Mirror Langevin diffusion under non-strongly log-concave target distributions. By constructing a Lyapunov function on the Hessian manifold and leveraging entropy-optimal transport together with strong data processing inequalities, the authors establish sufficient conditions guaranteeing the validity of Poincaré or log-Sobolev inequalities. The primary contribution lies in the introduction of a theoretically grounded two-step Gibbs sampler as a Markov chain approximation to the continuous diffusion process. The paper proves that this discrete-time sampler converges exponentially fast in χ² divergence at a rate matching that of the continuous dynamics, thereby providing the first rigorous convergence guarantee for such an algorithmic scheme in this setting.
📝 Abstract
Given a strongly convex function $u$, equip $R^d$ with a Riemannian metric given by the Hessian $\nabla^2 u$. This is a so-called Hessian manifold. Given a probability density $μ$ one may run a Langevin diffusion intrinsic to the manifold with stationary distribution $μ$. Such (Hessian) manifold-valued Langevin diffusions are called Mirror Langevin diffusions (MLD) which have recently become popular. One of the questions we explore is whether, given $μ$, one can choose $u$ to get an exponential convergence to equilibrium for the MLD, especially if $μ$ is not strongly log-concave. Our results are based on Lyapunov function methods and give sufficient conditions for a Poincaré or a log-Sobolev inequality to hold for the MLD. These, in turn, imply exponential convergence. We also introduce a Markov chain approximation to the MLD given by a two step Gibbs sampler with stationary distribution $μ$. This Markov chain is a variant of the Sinkhorn Markov chain introduced in arXiv:2307.16421 that is conjectured to converge to a time-inhomogeneous generalization of the MLD. Under suitable assumptions, we prove that the Markov chain has a guaranteed convergence rate in $χ^2$ that is consistent with the diffusion time scale. Our proofs are based on ideas from entropic optimal transport and strong data processing inequalities.