🤖 AI Summary
This study investigates the mechanism by which stochastic gradient descent (SGD) breaks scale symmetry in undercomplete linear autoencoders. By integrating principal component analysis (PCA) on manifolds with Hessian spectral theory, we reveal how loss geometry transforms gradient noise into directional drift along the manifold and establish an effective dynamical description of this process. Our findings demonstrate that such directed drift, driven by the geometry of the solution space, causes decoder weights to grow persistently toward the stability boundary, ultimately converging to solutions sharper than the equilibrium baseline. Furthermore, this work elucidates potential contradictions among different sharpness metrics.
📝 Abstract
Neural network loss landscapes have many symmetries, which are preserved by gradient flow but broken by finite-stepsize stochastic gradient descent (SGD). A canonical example of such a symmetry is scale in homogeneous networks: one can scale up the parameters in one layer and down in the next without changing the network output. Previous work has documented cases in which SGD breaks this symmetry in favor of balancing gradient noise or minimizing fluctuations. Here, we show that the solution geometry of undercomplete linear autoencoders instead selects a preferred sign for scale drift: on the PCA solution manifold, SGD favors large decoder weights. This directed scale drift occurs on a slow timescale, and its dynamics admit an analytically-tractable effective description. However, it cannot continue indefinitely: increasing scale eventually drives the dynamics towards a finite-stepsize stability boundary. The resulting solutions are sharper than a balanced baseline in the sense of the maximum eigenvalue of the loss Hessian, but different sharpness measures can move in opposing directions. Thus, undercomplete autoencoders give a concrete illustration of how loss geometry can convert residual gradient noise into directed motion along a manifold of functionally-equivalent solutions.