Exact information accounting for SGD methods

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of classical geometric analysis in precisely characterizing the convergence and generalization mechanisms of stochastic gradient descent (SGD) and its variants. By adopting an information-theoretic framework, this work models preconditioning steps as Gaussian Bayesian posterior mean updates and derives a unified information-theoretic identity encompassing convex convergence, saddle-point escape, and multiple SGD variants. This formulation elucidates the origins of looseness in classical bounds and reveals the fundamental distinctions among optimizers. Empirical evaluations on real-world neural networks quantify key metrics, effectively differentiating optimizers that achieve identical training losses. Furthermore, the analysis highlights the inherent limitations of generalization certificates derived under isotropic priors.
📝 Abstract
As an alternative to the standard geometric analyses, we give an exact, information-theoretic analysis of stochastic gradient descent (SGD) and its variants. We show that a preconditioned SGD step is the posterior-mean update of a Gaussian Bayes model, and that its one-step regret splits into an intrinsic-time cost and a change in comparator information. The split extends to an identity for the objective itself. Convex convergence, strict-saddle-point escape, the link between flatness and generalization, the standard learning-rate schedules, adaptive optimizers, and the noisy, momentum, heavy-tailed, and gradient-free variants of SGD each correspond to a term or a special case of this identity. We measure its terms on synthetic and real training runs. On real networks it attributes the slack of classical convergence bounds to the terms their derivations drop and separates optimizers that reach the same training loss. That separation follows the number and consistency of their steps. Its relation to which of them generalizes better differs between networks. For gradient-free SGD the identity determines how a curvature preconditioner should enter the update. The sharpness-based generalization certificate it yields, with a data-independent isotropic prior, is vacuous at network scale unless the curvature spectrum is nearly flat across all parameters.
Problem

Research questions and friction points this paper is trying to address.

Stochastic Gradient Descent
Information Theory
Generalization
Convergence Analysis
Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stochastic Gradient Descent
Information-Theoretic Analysis
Regret Decomposition
Gaussian Bayes Model
Generalization
🔎 Similar Papers
No similar papers found.