Optimizer Geometry Sets the Pace: Spectral Learning Dynamics in Matrix Factorization

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how optimizer geometry influences spectral bias and the emergence of low-rank solutions in matrix factorization. To this end, we construct a unified dynamical framework that precisely quantifies how normalization, curvature correction, and damping regulate singular value dynamics. By theoretically analyzing algorithms including Euclidean gradient descent, SignGD, Adam approximations, Muon, Shampoo, and K-FAC, we systematically reveal how optimizer geometry and network depth jointly govern training dynamics. Our analysis elucidates the intrinsic mechanisms through which different optimizers either eliminate or preserve low-rank bias within finite time. These findings provide rigorous theoretical foundations for designing optimizers tailored to specific inductive biases.
📝 Abstract
Recent successes of matrix- and curvature-based optimizers have renewed interest in how update geometry shapes learning. These methods normalize or precondition updates, changing how different components progress during training. In deep matrix factorization, the geometry that slows gradient descent (GD) favors low-rank solutions by delaying the emergence of small singular modes. This raises the question of what remains of that spectral bias when normalization or curvature correction weakens or removes the slowdown. We study how optimizer geometry and depth jointly govern this behavior through a common framework for singular mode dynamics. Under explicit balance and alignment assumptions, we derive birth, saturation, and decay laws for Euclidean GD, coordinate-wise updates of SignGD and an instantaneous Adam approximation, spectral updates of Muon and cumulative Shampoo, and a block curvature model of K-FAC. The resulting picture is not a simple ordering from stronger to weaker low-rank bias. To highlight, SignGD and ideal Muon eliminate the divergent birth barrier and drive unsupported modes to zero in finite time. Cumulative Shampoo initially retains GD's depth-dependent barrier, then accumulated gradients produce a catch-up phase while making previously active modes increasingly persistent. Undamped K-FAC cancels the factorization-induced slowdown while preserving the ordering of the target singular values, whereas positive damping introduces a spectral threshold below which the slow GD phase laws reappear. These results give normalization, accumulated state, damping, and depth a direct interpretation as controls determining when modes emerge, persist, and disappear during training.
Problem

Research questions and friction points this paper is trying to address.

matrix factorization
optimizer geometry
spectral learning dynamics
low-rank bias
singular mode dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Matrix Factorization
Optimizer Geometry
Spectral Dynamics
Low-rank Bias
Curvature Preconditioning
🔎 Similar Papers
No similar papers found.