Optimizer-dependent training dynamics converge to the same one-third optimal data scaling

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study clarifies the conflation between dynamical exponents and optimal data exponents in neural scaling laws under adaptive optimizers. By proposing a radial-tangential decomposition model for parameter updates, combined with an online teacher-student framework and stochastic dynamics analysis, we reveal a unified dynamical relation across optimizers: 2α_r + α_t = 1. Experimental validation on seven optimizers, including SGD, Adam, and Muon, demonstrates that although optimizers alter single-step learning rates, the tuned optimal loss envelopes consistently follow a D^{-1/3} power-law scaling. This work theoretically elucidates the intrinsic mechanism underlying the universality of the data exponent.
📝 Abstract
Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a $1/3$ exponent emerges from learning peaked distributions. That account describes SGD, but models in practice are trained with adaptive optimizers. Here we separate two exponents the $1/3$ account does not distinguish: how fast the loss falls with training steps along a single run, and how fast the optimally tuned loss falls with dataset size $D$. We show that the first, a dynamic exponent, is optimizer-specific while the second, an optimal data exponent, converges to $1/3$ across optimizers. In an online teacher-student model we decompose the loss into norm growth (radial) and alignment toward the teacher direction (tangential), each decaying as a power law with dynamic exponents $α_{r}$ and $α_{t}$. Under SGD, both are close to $1/3$, so the data exponent is also $1/3$ across different learning rates. Under Adam the two separate: $α_{r} \simeq 0.48$ but $α_{t} \simeq 0.08$. Since the total loss is minimized when these two parts are balanced, the optimal learning rate is optimizer-dependent: $D$-independent for SGD but falls with $D$ for Adam. Yet tuned to that optimum, the loss returns to $D^{-1/3}$ for both. A stochastic-dynamics analysis explains why: the optimizers can trade decay speed between the two channels, but they all fall on a single dynamic exponent relation, $2α_{r}+ α_{t} = 1$, which fixes the optimal data exponent at $1/3$. Across seven optimizers, including Muon, the measured exponents are consistent with this relation, and the optimal-loss envelopes agree with $D^{-1/3}$ across them. The optimizer sets how fast a model learns per step; tuned optimally, it changes the prefactor but not the rate at which loss falls per sample.
Problem

Research questions and friction points this paper is trying to address.

neural scaling laws
optimizer-dependent dynamics
data scaling exponent
adaptive optimizers
training dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neural scaling laws
Optimizer-dependent dynamics
Teacher-student model
Stochastic dynamics
Data scaling exponent
🔎 Similar Papers