🤖 AI Summary
This study addresses the deviation of neural network training from classical gradient flow under large learning rates, which hinders sparse solution recovery. Focusing on diagonal linear networks, this work introduces the concept of "drift" and reveals its competitive interplay with "gain," another form of implicit bias. Through theoretical analysis and targeted intervention strategies, we demonstrate that actively modulating gain can effectively guide model selection. Our findings indicate that large learning rates do not universally impede sparse recovery; rather, by controlling the competition between these implicit biases, effective optimization becomes achievable, enabling the recovery of sharper and sparser solutions.
📝 Abstract
Large learning rates can qualitatively change the trajectory of neural network training, often pushing optimization into regimes far from classical gradient-flow behavior. The Edge of Stability (EoS) offers a valuable lens on the dynamics such learning rates induce. We study corresponding dynamics in diagonal linear networks, where we uncover a competition between two distinct implicit biases that jointly determine the sparsity of the recovered solution in regression settings. Complementary to the Gain, which captures the average discretization error accumulated by Gradient Descent relative to Gradient Flow, we derive a closely associated but overlooked quantity: the Drift. Under large learning rates, it describes an imbalance between different discretization errors and represents a systematic shift in the optimization trajectory. While the Gain grows monotonically in certain regimes, and can bias towards denser, flatter interpolators, the impact of the Drift depends on its alignment with potential solutions, which can either counteract or reinforce the effect of the Gain. Consequently, its behavior drives model selection, particularly during early training epochs. To validate our theoretical insights, we introduce an intervention that actively steers the Gain to recover sharper, sparser solutions. Thus, our analysis reveals that large learning rates do not universally hinder the recovery of sparse solutions. On the contrary, they can be harnessed to control the implicit bias of training.