The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge that the Adam optimizer’s coupled moment estimation renders its adaptive mechanism intractable and incurs substantial memory overhead. By revealing a hidden scale-stable ratio under the Tied-Ξ² setting, this work reconstructs the algorithm to compress second-moment states while elucidating its theoretical connection to sign-based methods. Furthermore, leveraging heavy-tailed distribution properties, it proposes a scaling-free quantization scheme based on a 4-bit codebook alongside a learning rate transfer rule. The resulting approach achieves efficient optimization using only 4-bit storage, delivering performance comparable to full-precision Adam while significantly reducing memory consumption and simplifying model deployment.
πŸ“ Abstract
Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-$\beta$ regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically, this transformed ratio exhibits a stable, heavy-tailed distribution across tasks, model scales, and training stages, in contrast to the variability of raw moment magnitudes. This empirical stability has both practical and conceptual consequences. First, we derive a recurrence for the transformed ratio, yielding a reparameterization of Adam that replaces the second moment with a compressible state. Leveraging its stable distribution, we show that a fixed 4-bit codebook is sufficient in our experiments to store this state without auxiliary scaling, achieving performance competitive with full-precision Adam. Second, the transformed ratio view clarifies Adam's connection to sign-based methods: Adam reduces to sign-based momentum modulated by the transformed ratio, and replacing it with a constant recovers Signum as a limiting case. This perspective further provides a simple rule for transferring learning rates between the two methods. Together, these results suggest that tied-$\beta$ Adam admits a simple and approximately stable ratio structure underlying its adaptive behavior and demonstrate its utility for both analysis and efficient implementation.
Problem

Research questions and friction points this paper is trying to address.

Adam optimizer
adaptive behavior
exponential moving averages
deep neural networks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adam optimizer
Transformed ratio
4-bit compression
Sign dynamics
Reparameterization
πŸ”Ž Similar Papers
No similar papers found.