🤖 AI Summary
This paper addresses the poor robustness of Markov chain Monte Carlo (MCMC) methods under two pathological target distributions: *roughness* (characterized by highly oscillatory gradients causing numerical instability) and *flatness* (marked by vanishing gradients hindering exploration). We propose a unified robust MCMC framework that formally characterizes both pathologies and integrates non-local jumps, adaptive step sizes, gradient regularization, and stability-driven proposal mechanisms. Theoretical analysis exposes the convergence failure mechanisms of conventional local MCMC algorithms under such pathologies. Empirical evaluation demonstrates substantial improvements in sampling efficiency and convergence speed on canonical pathological benchmarks. Key contributions include: (i) establishing quantitative, interpretable criteria for diagnosing roughness and flatness; (ii) systematically unifying anti-pathological design principles into a coherent framework; and (iii) providing a theoretically grounded yet practical sampling methodology applicable to high-dimensional, nonsmooth, and low signal-to-noise ratio settings.
📝 Abstract
Markov Chain Monte Carlo (MCMC) is a flexible approach to approximate sampling from intractable probability distributions, with a rich theoretical foundation and comprising a wealth of exemplar algorithms. While the qualitative correctness of MCMC algorithms is often easy to ensure, their practical efficiency is contingent on the `target' distribution being reasonably well-behaved.
In this work, we concern ourself with the scenario in which this good behaviour is called into question, reviewing an emerging line of work on `robust' MCMC algorithms which can perform acceptably even in the face of certain pathologies.
We focus on two particular pathologies which, while simple, can already have dramatic effects on standard `local' algorithms. The first is roughness, whereby the target distribution varies so rapidly that the numerical stability of the algorithm is tenuous. The second is flatness, whereby the landscape of the target distribution is instead so barren and uninformative that one becomes lost in uninteresting parts of the state space. In each case, we formulate the pathology in concrete terms, review a range of proposed algorithmic remedies to the pathology, and outline promising directions for future research.