π€ AI Summary
This study addresses the absence of global convergence guarantees for third-order Langevin dynamics in non-convex optimization. Within a simulated annealing framework, it proposes an explicit three-block distorted entropy method, establishing global convergence rates under logarithmic cooling by combining entropy transfer with dissipation assumptions. By employing exact force integration and midpoint three-stage discretization, the work derives more relaxed step-size conditions, controlling local errors through strong coupling analysis and polynomial decay strategies. Theoretically, it establishes kinetic energy-controlled convergence rates. Empirically, experiments on double-well potentials and neural network tasks demonstrate that the proposed approach outperforms overdamped dynamics, significantly improving convergence success rates.
π Abstract
We study global convergence guarantees of third-order Langevin dynamics for non-convex optimization via simulated annealing with fixed friction and decreasing noise. An explicit three-block distorted entropy transfers dissipation from the noisy auxiliary variable to the full state. Under dissipativity, regularity, and low-temperature functional-inequality assumptions, logarithmic cooling drives the objective values to the global minimum in probability at the barrier-controlled kinetic rate. For the exact-force-integral and midpoint three-stage discretizations, polynomially decreasing steps preserve this rate on the physical time scale. The cubic local endpoint estimate gives a less restrictive sufficient step-size condition than the available frozen-force kinetic result. A comparison with the one-gradient UBU integrator shows how its centered stochastic local error leads, under the same strong-coupling analysis, to a smaller sufficient iteration exponent. Numerical experiments are conducted to illustrate our theory. For a double well objective, third-order Langevin terminal-success point estimates are higher than UBU at both a common horizon and an equal gradient budget. For a high-dimensional nonconvex neural-network objective using synthetic data, independently tuned UBU and third-order Langevin schemes both outperform overdamped Langevin dynamics; the third-order Langevin point estimate is higher. For the same neural-network objective on real data, we show the same point-estimate ordering for best-basin probability and post-quench test accuracy. Numerical code and associated experiment results are publicly available at https://github.com/gagawjbytw/simulated-annealing-third-order-langevin.