High-Probability Guarantees for SGD under $\beta$-Heavy-Tailed Gradient Noise

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the failure of high-probability guarantees for stochastic gradient descent (SGD) caused by the heavy-tailed nature of deep learning gradient noise. Moving beyond traditional sub-Gaussian assumptions, it employs Young functions in Orlicz spaces to uniformly model $\beta$-heavy-tailed noise, establishing a theoretical framework encompassing heavier distributions such as the log-normal. By integrating concentration inequalities, uniform trajectory bounds, and the Polyak-Łojasiewicz condition, this work derives high-probability optimization and generalization risk bounds for SGD under non-convex losses, while systematically analyzing the role of gradient clipping. Ultimately, it obtains high-probability convergence and last-iterate risk bounds that do not rely on specific learning rate decay schedules, explicitly characterizing the interplay between noise tail behavior and learning rates.
πŸ“ Abstract
Stochastic gradient descent (SGD) is widely used to train machine learning models, but subsampling the training data introduces noise into its updates. The strength and applicability of high-probability guarantees therefore depend critically on how the tails of gradient noise are modeled. Reports of heavy-tailed gradient noise in deep learning motivate relaxing the bounded-noise and sub-Gaussian assumptions commonly used in high-probability analyses of SGD. We use Young functions from Orlicz space theory to describe noise tails in a common framework. We model SGD gradient noise by adopting a Young function that preserves the finiteness of all polynomial moments while allowing tails heavier than sub-Weibull, including lognormal distributions. The resulting class is called $\beta$-heavy-tailed, with $\beta$ controlling the tail heaviness. We establish concentration inequalities for $\beta$-heavy-tailed noise and combine them with a uniform bound on the difference between empirical and population gradients along the SGD trajectory to obtain high-probability bounds on optimization and population-risk stationarity for smooth nonconvex losses under trajectory assumptions. The bounds are not restricted to a particular learning-rate decay rule and make explicit the effects of noise tails and learning-rate schedules. Under the Polyak-{\L}ojasiewicz condition, we bound the risk at the last iterate. We also analyze SGD with gradient clipping under the $\beta$-heavy-tailed noise model.
Problem

Research questions and friction points this paper is trying to address.

Stochastic Gradient Descent
Heavy-Tailed Noise
High-Probability Guarantees
Nonconvex Optimization
Concentration Inequalities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stochastic Gradient Descent
Heavy-Tailed Noise
Orlicz Space
High-Probability Guarantees
Gradient Clipping