π€ AI Summary
This study addresses the failure of high-probability guarantees for stochastic gradient descent (SGD) caused by the heavy-tailed nature of deep learning gradient noise. Moving beyond traditional sub-Gaussian assumptions, it employs Young functions in Orlicz spaces to uniformly model $\beta$-heavy-tailed noise, establishing a theoretical framework encompassing heavier distributions such as the log-normal. By integrating concentration inequalities, uniform trajectory bounds, and the Polyak-Εojasiewicz condition, this work derives high-probability optimization and generalization risk bounds for SGD under non-convex losses, while systematically analyzing the role of gradient clipping. Ultimately, it obtains high-probability convergence and last-iterate risk bounds that do not rely on specific learning rate decay schedules, explicitly characterizing the interplay between noise tail behavior and learning rates.
π Abstract
Stochastic gradient descent (SGD) is widely used to train machine learning models, but subsampling the training data introduces noise into its updates. The strength and applicability of high-probability guarantees therefore depend critically on how the tails of gradient noise are modeled. Reports of heavy-tailed gradient noise in deep learning motivate relaxing the bounded-noise and sub-Gaussian assumptions commonly used in high-probability analyses of SGD. We use Young functions from Orlicz space theory to describe noise tails in a common framework. We model SGD gradient noise by adopting a Young function that preserves the finiteness of all polynomial moments while allowing tails heavier than sub-Weibull, including lognormal distributions. The resulting class is called $\beta$-heavy-tailed, with $\beta$ controlling the tail heaviness. We establish concentration inequalities for $\beta$-heavy-tailed noise and combine them with a uniform bound on the difference between empirical and population gradients along the SGD trajectory to obtain high-probability bounds on optimization and population-risk stationarity for smooth nonconvex losses under trajectory assumptions. The bounds are not restricted to a particular learning-rate decay rule and make explicit the effects of noise tails and learning-rate schedules. Under the Polyak-{\L}ojasiewicz condition, we bound the risk at the last iterate. We also analyze SGD with gradient clipping under the $\beta$-heavy-tailed noise model.