Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima

📅 2026-07-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the theoretical gap in large-stepsize gradient descent for over-parameterized least squares problems with vector outputs, where conventional theory requires stepsizes smaller than twice the inverse sharpness—yet practitioners routinely use larger stepsizes near flat minima manifolds without rigorous justification. By integrating tools from dynamical systems, differential geometry, and singular partial differential equations, this study establishes the first convergence theory for large-stepsize gradient descent extended from isolated flat minima to general flat minima manifolds. The main contributions include a unified normal form and convergence analysis, the discovery that the set of minimizers in deep matrix factorization possesses a fiber bundle structure over a product of spheres satisfying the Morse–Bott condition, and a novel method for solving the associated singular PDEs, thereby providing a rigorous theoretical foundation for large-stepsize optimization.
📝 Abstract
An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian. Classical analyses typically require the step size to be uniformly smaller than twice the reciprocal of the sharpness, but this condition is frequently violated in the training of deep neural networks. Recent work bridges this gap in the setting of overparametrised least-squares with a \emph{single scalar output}, providing a normal form for large-step GD in a neighbourhood of an \emph{isolated} flat minimum and establishing three corresponding convergence results. In this paper, we extend this theory in two directions: (1) to overparametrised least-squares with \emph{vector-valued outputs} (including regression with arbitrarily many observations), and (2) to a neighbourhood of a \emph{manifold} of flat minima (which we show is essential for applications such as matrix factorisation). We generalise both the normal form and all three convergence theorems of \cite{macdonaldeos} to this broader setting, overcoming several technical challenges, including the solution of a singular partial differential equation via a novel method that may be of independent interest. We further show that our framework applies to deep matrix factorisation under mild assumptions, yielding several new structural results. In particular, we prove that the set of flat minima forms a fibre bundle over a product of spheres, and that the sharpness is Morse-Bott along this manifold.
Problem

Research questions and friction points this paper is trying to address.

gradient descent
large step size
flat minima manifold
overparametrised least-squares
sharpness
Innovation

Methods, ideas, or system contributions that make the work stand out.

large-step gradient descent
manifold of flat minima
overparametrised least-squares
Morse-Bott sharpness
singular PDE
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lachlan Ewen MacDonald
Innovation in Data Engineering and Science (IDEAS), University of Pennsylvania, Pennsylvania, PA 19104
R
René Vidal
Innovation in Data Engineering and Science (IDEAS), University of Pennsylvania, Pennsylvania, PA 19104