π€ AI Summary
This study addresses the challenge that the Adam optimizerβs coupled moment estimation renders its adaptive mechanism intractable and incurs substantial memory overhead. By revealing a hidden scale-stable ratio under the Tied-Ξ² setting, this work reconstructs the algorithm to compress second-moment states while elucidating its theoretical connection to sign-based methods. Furthermore, leveraging heavy-tailed distribution properties, it proposes a scaling-free quantization scheme based on a 4-bit codebook alongside a learning rate transfer rule. The resulting approach achieves efficient optimization using only 4-bit storage, delivering performance comparable to full-precision Adam while significantly reducing memory consumption and simplifying model deployment.
π Abstract
Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-$\beta$ regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically, this transformed ratio exhibits a stable, heavy-tailed distribution across tasks, model scales, and training stages, in contrast to the variability of raw moment magnitudes. This empirical stability has both practical and conceptual consequences. First, we derive a recurrence for the transformed ratio, yielding a reparameterization of Adam that replaces the second moment with a compressible state. Leveraging its stable distribution, we show that a fixed 4-bit codebook is sufficient in our experiments to store this state without auxiliary scaling, achieving performance competitive with full-precision Adam. Second, the transformed ratio view clarifies Adam's connection to sign-based methods: Adam reduces to sign-based momentum modulated by the transformed ratio, and replacing it with a constant recovers Signum as a limiting case. This perspective further provides a simple rule for transferring learning rates between the two methods. Together, these results suggest that tied-$\beta$ Adam admits a simple and approximately stable ratio structure underlying its adaptive behavior and demonstrate its utility for both analysis and efficient implementation.