The Devil Is in the Reconstruction Loss Scale: Rethinking Optimization in LLM Quantization

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the optimization imbalance in post-training quantization of large language models, where mean squared error (MSE) induces disproportionate gradients across quantization stages due to scale disparities in reconstruction losses. Grounded in deep learning optimization theory, this work proposes a principle that decouples optimization intensity from loss magnitude and introduces a novel paradigm leveraging root mean squared error (RMSE) to achieve implicit gradient normalization. The contributions reveal the underlying mechanism by which loss scale influences quantization optimization and demonstrate that RMSE serves as an effective alternative to MSE. By significantly balancing optimization intensity across quantization stages, the proposed approach enhances overall model performance.
📝 Abstract
Post-training quantization (PTQ) methods typically use sequential quantization that partitions a pre-trained LLM into a series of units (e.g., transformer blocks), with one unit quantized at each stage. State-of-the-art PTQ methods are predominantly learning-based, optimizing auxiliary quantization parameters (e.g., scaling factors, rotation matrices, clipping thresholds, and adapters) via gradient descent to minimize a reconstruction loss. A common practice is to use mean squared error (MSE) as the reconstruction loss function, yet its induced optimization behavior remains largely unexplored. In this work, we take a holistic view of sequential quantization and systematically investigate how optimization evolves from the first quantization stage to the last, aiming for a deep understanding of optimization in learning-based PTQ schemes. Through extensive empirical studies spanning representative learning-based PTQ methods, LLM families, model scales, architectures, quantization settings and various tasks, we consistently uncover Optimization Imbalance: reconstruction loss magnitudes vary dramatically across stages, accompanied by highly uneven gradient magnitudes and parameter updates under MSE. We term the cross-stage range of loss magnitudes the reconstruction loss scale, and reveal that MSE translates the unexpectedly large reconstruction loss scale into highly uneven gradient magnitudes, which in turn lead to uneven optimization strength across quantization stages. This finding suggests a general principle for improving learning-based PTQ: optimization strength across stages should be decoupled from the reconstruction loss scale. Theoretically, we show that root mean squared error (RMSE) variants defined at the sample, channel, token, and element levels naturally realize this principle through implicit gradient normalization, outperforming MSE significantly as a drop-in replacement.
Problem

Research questions and friction points this paper is trying to address.

Post-training quantization
Large Language Models
Reconstruction loss
Optimization imbalance
Mean squared error
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-training quantization
Reconstruction loss scale
Optimization imbalance
Root mean squared error
Implicit gradient normalization
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Chao Li
Chao Li
Intel Labs China
S
Shigeng Wang
Intel Labs China
A
Anbang Yao
Intel Labs China