Rare Gate Disagreements Can Limit Plasticity: When Gradient Flow Mispredicts Finite-Batch SGD

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the qualitative failure of population gradient flow in predicting finite-batch SGD adaptivity, which obscures the mechanisms underlying limited network plasticity after pretraining. By leveraging rigorous theoretical derivations and numerical simulations on ReLU regression models, this work abandons diffusion approximations to analyze cumulative probabilities and conditional expectation contraction within online recursions. It provides the first precise quantification of the relationship between gating divergence probability under weight decay and recovery time, revealing that source-task-induced neuronal angle contraction necessitates exponentially many samples for target recovery. Furthermore, it demonstrates that the failure of small-step-size SGD following prolonged pretraining stems from joint limit effects, with recovery performance constrained by a divergence budget and saturating within specific horizons.
📝 Abstract
Population gradient flow is a common tool for reasoning about how neural networks adapt, including after pretraining. We show that it can mispredict finite-batch stochastic gradient descent (SGD) qualitatively, and we trace the discrepancy to a specific mechanism. In a two-unit ReLU regression, a source task drives the two neurons toward positive proportionality and a target task rewards separating them. After source training for time $T$, gradient flow recovers on the target in time linear in $T$. Online SGD with batch size $b$ and step size $η$ in both phases instead fails with high probability throughout a horizon of order $e^{c/η}$ once $T \gtrsim \log(b/η)$, uniformly on an explicit set of initializations with Gaussian probability above one percent. For each fixed $T$, small-step SGD still recovers, so the failure requires the joint limit of small steps and long pretraining. At the target clone, the population instability is carried entirely by inputs on which the two ReLU gates disagree. For units at angle $δ$ these inputs form a wedge of probability $δ/π$, and weight decay shrinks the angle exponentially during pretraining. On every other input both units receive the same random linear update, which contracts their separation in conditional expectation. Bounding the cumulative probability of sampling the wedge along the exact online recursion, without a diffusion approximation, shows that recovery with fixed probability from an identical source-gradient-flow checkpoint, within $e^{c/η}$ updates, requires $Nb \gtrsim e^{λT}$ target samples and batch size $b \gtrsim ηe^{λT}$, where $N$ counts updates and $λ$ is the weight decay. In simulations, recovery is approximately a function of the disagreement budget $bδ/η$ and saturates in the horizon.
Problem

Research questions and friction points this paper is trying to address.

gradient flow
finite-batch SGD
plasticity
gate disagreements
pretraining
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gradient Flow
Finite-Batch SGD
Plasticity
Gate Disagreements
Weight Decay
Ruoyu Zhao
Ruoyu Zhao
Jiangxi University of Finance and Economics
Multimedia SecurityPrivacyUsable Security and Privacy
M
Mingxuan Zhang
Microsoft
J
Jianbo Dai
Copula Lab
J
Jiaqi Wu
City University of Hong Kong
C
Chenyu Zhu
City University of Hong Kong
T
Tong Che
NVIDIA Research