Learning the identity: a case study of how SGD selects among functional decompositions

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the question of why stochastic gradient descent (SGD) favors specific function decompositions in identity mapping learning, a phenomenon that existing theories struggle to explain regarding its implicit bias among equivalent solutions. To this end, this work proposes a theoretical framework grounded in deep linear residual networks, introducing an entropy loss perspective to quantitatively analyze SGD dynamics on non-convex optimization manifolds and reveal how parameterization substantively shapes optimization trajectories. Theoretical predictions align closely with empirical results, elucidating the intrinsic mechanisms by which SGD selects particular function decompositions. Ultimately, this research establishes a novel paradigm for understanding the implicit regularization effects of optimization algorithms in deep learning.
📝 Abstract
One might think that learning the identity function with a deep linear residual network is trivial - the path along residual connections already implements the identity, and so the network need only drive its weights to zero. However, this zero-weight solution is just one point on an entire manifold of population-loss minimizers, each corresponding to a different decomposition of the identity across the network's layers. Although the population loss does not distinguish among these solutions, stochastic gradient descent (SGD) reproducibly favors particular ones. For instance, under anisotropic label noise, the learned layers exhibit a noise-dependent spectrum; even with weight decay, SGD does not generally recover the zero-weight solution. Changing only the parametrization, while leaving the set of realizable functions unchanged, yields different behavior: factoring each weight matrix as a product of two matrices causes the weights to collapse to zero, even without explicit weight decay. While perhaps mysterious and unintuitive at first, these phenomena can be understood through the lens of entropic loss, which augments the population loss with a term proportional to the expected squared norm of the minibatch gradient (Ziyin et al., 2025). On the identity manifold, the population loss is constant, while the entropic term distinguishes among these decompositions. We characterize its minimizers analytically and use them to derive predictions for the structure of solutions favored by SGD. Networks trained with SGD closely match these predictions. Overall, the identity learning task studied here serves as a clean and simple case study of how the lens of entropic loss can clarify why SGD favors particular decompositions of the same input-output function.
Problem

Research questions and friction points this paper is trying to address.

stochastic gradient descent
identity learning
functional decomposition
entropic loss
deep linear residual network
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stochastic Gradient Descent
Entropic Loss
Functional Decomposition
Deep Linear Residual Networks
Identity Learning
🔎 Similar Papers
2024-05-25arXiv.orgCitations: 3