🤖 AI Summary
Standard weight decay neglects the spectral structure of weight matrices, limiting its ability to effectively induce low-rank properties. This work proposes Spectral Weight Decay, which achieves additive spectral shrinkage through nuclear norm regularization and decoupled optimization. We establish a theoretical connection between this approach and approximate proximal descent, overcoming the limitations of conventional L2 regularization in promoting low rank. Leveraging singular value decomposition, the method is applied to large-scale models such as LLaMA and BERT. Experiments demonstrate that on a 500M-parameter model, it achieves a 1.89× compression ratio and a 1.18× inference speedup. Furthermore, accuracy improves by up to 17.8% in high-noise settings, significantly enhancing both model compression efficiency and robustness.
📝 Abstract
Standard weight decay treats each weight matrix as a vector and ignores its spectral structure. We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage. We connect the update to approximate proximal descent and show that its sensitivity to update order can exceed that of conventional $\ell_2$ weight decay near rank deficiency. Across LLaMA models with $124$M to $500$M parameters, spectral weight decay lowers effective rank and improves SVD-LLM compression at matched validation loss. At $500$M and a $4\%$ distortion budget, it reaches $1.89\times$ compression and $1.18\times$ GPU inference speedup, compared with $1.14\times$ and $1.01\times$ after standard weight decay. Under fixed-horizon training with $60\%$ label noise, it also improves final mean clean-test accuracy over matched $\ell_2$ regularization by up to $17.8$ points on MNIST and $4.6$ points across four BERT-base tasks. Code is available at https://github.com/brain-lab-research/SpectralWD.