When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the paradox in low-bit quantization whereby minimizing reconstruction loss often leads to performance degradation under distribution shifts, indicating that low reconstruction error may compromise generalization. To mitigate this, it proposes a posterior distributionally robust refinement framework that requires no modifications to inference operators. Grounded in distributionally robust optimization theory, the approach constrains an ambiguity set of input activation distributions and fine-tunes integer codes within existing quantization grids to optimize weights by minimizing worst-case reconstruction loss. The proposed method significantly enhances the downstream task performance of six mainstream quantization techniques, including AWQ and GPTQ. Furthermore, it is compatible with both dense and Mixture-of-Experts (MoE) architectures without introducing additional inference overhead, thereby effectively strengthening the generalization capability of large language models.
📝 Abstract
Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision. We show that the weights favored by minimizing this loss need not yield better model performance on new tasks. In fact, we find that lower reconstruction loss can even degrade model performance on the same calibration data. Our analysis further shows that weights with lower reconstruction loss on calibration data can have higher loss than other weights when the distribution of input activations changes. Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions. DRQ refines the integer codes representing quantized weights within the existing quantization grid, keeping quantization parameters and inference operators unchanged. Extensive experiments show that DRQ improves models quantized by six representative PTQ methods, including AWQ, GPTQ, and ParoQuant, and delivers gains across both dense and mixture-of-experts large language models. These results establish DRQ as a general post-hoc refinement framework for weight-only PTQ, achieving better downstream performance without adding inference overhead.
Problem

Research questions and friction points this paper is trying to address.

Post-Training Quantization
Reconstruction Loss
Low-Bit Quantization
Large Language Models
Distribution Shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distributionally Robust Quantization
Post-Training Quantization
Reconstruction Loss
Low-Bit LLM Quantization
Post-hoc Refinement
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yanlong Zhao
University of Science and Technology of China
X
Xiaoyuan Cheng
University College London
H
Huihang Liu
Shanghai University of Finance and Economics
B
Baihua He
University of Science and Technology of China
X
Xinyu Zhang
University of Science and Technology of China, AMSS, Chinese Academy of Sciences
Harrison Bo Hua Zhu
Harrison Bo Hua Zhu
Assistant Professor, University of Copenhagen
EpidemiologyPhylogeneticsInfectious DiseasesProbabilistic Machine LearningDeep Learning
Wenlong Chen
Wenlong Chen
Research Scientist, Isomorphic Labs
Machine LearningDeep LearningArtificial Intelligence
Li Zeng
Li Zeng
Peking University
LLM training and inferenceVector ComputingGraph Computing
Zhuo Sun
Zhuo Sun
Australian National University
Wireless Comunications