🤖 AI Summary
This work proposes SG-TULA, a subgradient-based explicit tamed unadjusted Langevin algorithm for sampling from target distributions defined by nonsmooth, nonconvex potential functions with superlinearly growing gradients, without requiring any smoothing preprocessing. It establishes, for the first time, an explicit non-asymptotic Wasserstein-2 convergence bound for subgradient Langevin methods under this challenging setting, precisely characterizing the dependence on dimension and temperature, and derives corresponding excess risk bounds for optimization. When applied to GPT-2 pretraining, SG-TULA not only outperforms AdamW and Muon in empirical performance but also provides non-asymptotic theoretical guarantees that are absent in existing methods.
📝 Abstract
We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoothing procedures. To handle the superlinear regime, taming techniques are employed to produce a stable, explicit scheme. We derive non-asymptotic convergence bounds in Wasserstein-2 distance, with all constants tracked explicitly in terms of dimension and inverse temperature, improving upon the currently known rates for subgradient-based Langevin algorithms. We further provide excess risk estimates for the associated optimisation problem. We verify the assumptions, with explicit constants, for the regularized pretraining potential of a LLM in the GPT-2 lineage and the boosted coordinate-wise variant of SG-TULA pretrains the former competitively against finetuned AdamW and Muon, for which no comparable non-asymptotic guarantees are presently available.