Clipped or Unclipped? Finite-Sample Trade-offs for Averaged SGD under Heavy-Tailed Noise

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unresolved question of whether gradient clipping genuinely improves the statistical accuracy of averaged stochastic gradient descent (SGD) under heavy-tailed noise. Under a bounded p-th moment assumption, the authors employ one-dimensional quadratic recursion techniques to derive finite-sample accuracy trade-offs between clipped and unclipped algorithms, establishing novel error bounds that balance bias and concentration. The work proves the tightness of existing unclipped bounds and reveals that, under higher-moment conditions, clipping does not necessarily enhance performance and may incur additional asymptotic variance costs. Furthermore, it reduces the failure probability dependence from polynomial to logarithmic order. These findings provide a rigorous theoretical foundation for understanding the mechanisms and limitations of gradient clipping in heavy-tailed optimization settings.
📝 Abstract
Gradient clipping is widely used to stabilize training, but it need not improve the statistical accuracy of averaged SGD, even under heavy-tailed noise. We derive a finite-sample comparison of clipped and unclipped Polyak-Ruppert averaged SGD under finite conditional $p$-th moments, $p\ge2$. Our main result gives explicit accuracy and confidence conditions under which, for $p>2$, the Gaussian term dominates the unclipped deviation bound, so clipping need not improve its leading order. By balancing clipping bias and concentration, we obtain a bound in which the heavy-tail correction depends logarithmically rather than polynomially on the inverse failure probability. At $p=2$, this improves the confidence dependence of the leading bound. We establish sharpness of the unclipped heavy-tail term through an exact one-dimensional quadratic recursion and extend the comparison to projected convex SGD. We also prove concrete costs of clipping: every fixed finite threshold increases asymptotic variance on a scalar Gaussian quadratic, while whole-gradient clipping can shift the limiting point under asymmetric noise.
Problem

Research questions and friction points this paper is trying to address.

gradient clipping
averaged SGD
heavy-tailed noise
finite-sample analysis
stochastic optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gradient Clipping
Averaged SGD
Heavy-Tailed Noise
Finite-Sample Analysis
Polyak-Ruppert Averaging
🔎 Similar Papers
No similar papers found.