Stable-SPAM: How to Train in 4-Bit More Stably than 16-Bit Adam

📅 2025-02-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the severe gradient norm fluctuations and training divergence under high learning rates in 4-bit low-precision large language model (LLM) training. To tackle these challenges, we propose Stable-SPAM, a stabilization-oriented optimizer. Its core innovations are: (1) an adaptive peak clipping mechanism based on historical gradient maxima; (2) full-matrix gradient normalization driven by sliding-window ℓ²-norm statistics; and (3) a spike-aware momentum resetting strategy with periodic reinitialization. Evaluated on 4-bit LLaMA-1B training, Stable-SPAM reduces gradient norm variance by 73% compared to BF16 Adam, achieves a 2.0 reduction in perplexity, cuts convergence steps by 50%, and—critically—demonstrates, for the first time, that 4-bit training can outperform its full-precision (BF16) baseline in both efficiency and final model quality.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Learning & Optimization for NLPSearch and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
This paper comprehensively evaluates several recently proposed optimizers for 4-bit training, revealing that low-bit precision amplifies sensitivity to learning rates and often causes unstable gradient norms, leading to divergence at higher learning rates. Among these, SPAM, a recent optimizer featuring momentum reset and spike-aware gradient clipping, achieves the best performance across various bit levels, but struggles to stabilize gradient norms, requiring careful learning rate tuning. To address these limitations, we propose Stable-SPAM, which incorporates enhanced gradient normalization and clipping techniques. In particular, Stable-SPAM (1) adaptively updates the clipping threshold for spiked gradients by tracking their historical maxima; (2) normalizes the entire gradient matrix based on its historical $l_2$-norm statistics; and $(3)$ inherits momentum reset from SPAM to periodically reset the first and second moments of Adam, mitigating the accumulation of spiked gradients. Extensive experiments show that Stable-SPAM effectively stabilizes gradient norms in 4-bit LLM training, delivering superior performance compared to Adam and SPAM. Notably, our 4-bit LLaMA-1B model trained with Stable-SPAM outperforms the BF16 LLaMA-1B trained with Adam by up to $2$ perplexity. Furthermore, when both models are trained in 4-bit, Stable-SPAM achieves the same loss as Adam while requiring only about half the training steps. Code is available at https://github.com/TianjinYellow/StableSPAM.git.
Problem

Research questions and friction points this paper is trying to address.

Enhances stability in 4-bit training
Reduces sensitivity to learning rates
Improves gradient norm stabilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Enhanced gradient normalization techniques
Adaptive clipping for spiked gradients
Momentum reset inherited from SPAM
Tianjin Huang
Tianjin Huang
Asst. Professor, CS@University of Exeter & Researcher Fellow, CS@TU/e
LLMsAdversarial examplesStable TrainingGraph Neural NetworkSparse Training
Haotian Hu
Haotian Hu
浙江大学
自动驾驶
Z
Zhenyu Zhang
Department of Electrical and Computer Engineering, University of Texas at Austin
Gaojie Jin
Gaojie Jin
Lecturer (Assistant Professor), University of Exeter
Machine LearningStatistical LearningTrustworthy AIHuman-GenAI-Alignment
X
Xiang Li
Department of Computer Science, University of Reading
L
Li Shen
School of Cyber Science and Technology, Sun Yat-sen University
Tianlong Chen
Tianlong Chen
Assistant Professor, CS@UNC Chapel Hill; Chief AI Scientist, hireEZ
Machine LearningAI4ScienceComputer VisionSparsity
L
Lu Liu
Department of Computer Science, University of Exeter
Q
Qingsong Wen
Squirrel Ai Learning
Z
Zhangyang Wang
Department of Electrical and Computer Engineering, University of Texas at Austin
S
Shiwei Liu
Mathematical Institute, University of Oxford