🤖 AI Summary
This study addresses the issue of weight oscillation during quantization-aware pretraining, which introduces noise and severely impedes convergence. To mitigate this, we propose CEWT, the first hyperparameter-free, zero-memory-overhead oscillation suppression technique. By applying a projection operation after optimizer updates, CEWT constrains the empirical weight distribution to a zero-mean Gaussian prior, effectively suppressing oscillations while seamlessly integrating with LLaMA and GPT architectures. This work significantly enhances low-precision training efficiency for large language models. Experimental results demonstrate that CEWT reduces perplexity by an average of 2.5 points, with reductions reaching up to 21 points, while incurring only 4% additional computational overhead. Furthermore, it supports extreme 1-bit quantized training.
📝 Abstract
Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise, they either introduce additional hyperparameters or memory overhead, or cannot consistently improve model accuracy. In this work, we propose optimization with $\textbf{C}$onstrained $\textbf{E}$mpirical $\textbf{W}$eight dis$\textbf{T}$ribution (CEWT), the first hyperparameter-free memory-overhead-free oscillation suppression method that consistently improves QAPT performance: an optimizer post-update step that projects weights to the nearest point in weight space whose empirical distribution (histogram) matches a zero-mean Gaussian. Our key insight is many quantizers are designed with the implicit assumption that the to-be-quantized data are permutations of samples from a zero-mean Gaussian, and this assumption is not true during QAPT. By enforcing the zero-mean Gaussian prior as a hard constraint, CEWT can suppress this detrimental noise. Empirical results on various combinations of SOTA quantizers and hypersphere optimizers suggest, that with a geomean increase of 4% in training time, CEWT can consistently reduce the pre-training perplexity (by an average of 2.5 and up to 21 points) of low-precision (down to 1-bit activations and weights and up to 610M parameters) LLaMA/GPT models without introducing any hyperparameters or storage overhead. Code is available at https://github.com/1733116199/cewt