Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of rounding errors accumulating recursively through moment estimation during optimizer state quantization in AdamW, which leads to preconditioner distortion and perturbed adaptive updates. To mitigate this, the work reformulates 4-bit AdamW quantization from a rounding-space perspective, proposing ZIP-SR (preconditioner-space stochastic rounding) and ZE-EDEN (zero-exclusion calibration) strategies, integrated with the NF4 format to optimize quantization dynamics. Evaluated on models ranging from 130M to 2.7B parameters, the proposed approach reduces the validation loss gap by 70%, achieving training performance closely approximating full 32-bit precision.
📝 Abstract
Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates. We redesign 4-bit optimizer-state quantization for AdamW from the perspective of \emph{rounding space}: the coordinate in which a quantizer chooses between adjacent reconstruction levels. For the second moment, a local analysis of the quantization cell adjacent to zero shows that small mean state error need not imply small mean preconditioner error at the next step. A one-dimensional quadratic construction further shows qualitatively different optimization dynamics under state-space and preconditioner-space rounding. These results motivate Zero-Inclusive Preconditioner-space Stochastic Rounding (\textbf{ZIP-SR}), which retains zero in the second-moment codebook and computes stochastic-rounding probabilities in preconditioner space. As a complementary route, Zero-Excluding EDEN calibration (\textbf{ZE-EDEN}) uses a zero-excluding second-moment codebook and rescales the quantized second-moment block to mitigate the preconditioner distortion caused by the positive quantization floor. Both configurations use 4-bit NormalFloat (NF4) for the first moment, with targeted stochastic rounding of the LM-head first moment during the final 10\% of training. Across GPT- and Llama-style pretraining experiments ranging from \textbf{130M} to \textbf{2.7B} parameters, both methods reduce TorchAO 4-bit AdamW's mean validation-loss gap to 32-bit AdamW at every evaluated model size, with the largest reported gap reduction reaching \textbf{70\%}. In full-parameter supervised fine-tuning, both recipes achieve lower validation loss than TorchAO while remaining close to 32-bit AdamW on downstream tasks.
Problem

Research questions and friction points this paper is trying to address.

AdamW optimizer-state quantization
preconditioner distortion
quantization error propagation
rounding space
4-bit quantization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Optimizer-State Quantization
Preconditioner-space Rounding
Stochastic Rounding
AdamW
4-bit NormalFloat
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
H
Hanyang Li
University of California, Berkeley
Shao Tang
Shao Tang
Linkedin
LLM Post-TrainingAgentOptimization
D
Daniel Thomas Braithwaite
Nubank
Gregory Dexter
Gregory Dexter
LinkedIn Corporation
Leonardo Neves
Leonardo Neves
Principal Research Engineer @ Snap Research and Teaching Fellow @ Harvard University
Recommender SystemsUser modelingNLPComputational Social Science
A
Aman Gupta
Nubank
H
Hiroto Udagawa
Nubank
A
Abhishek Shivanna
Nubank
D
Daniel Silva
Nubank
R
Rohan Ramanath
Nubank