🤖 AI Summary
This study addresses the paradox in low-bit quantization whereby minimizing reconstruction loss often leads to performance degradation under distribution shifts, indicating that low reconstruction error may compromise generalization. To mitigate this, it proposes a posterior distributionally robust refinement framework that requires no modifications to inference operators. Grounded in distributionally robust optimization theory, the approach constrains an ambiguity set of input activation distributions and fine-tunes integer codes within existing quantization grids to optimize weights by minimizing worst-case reconstruction loss. The proposed method significantly enhances the downstream task performance of six mainstream quantization techniques, including AWQ and GPTQ. Furthermore, it is compatible with both dense and Mixture-of-Experts (MoE) architectures without introducing additional inference overhead, thereby effectively strengthening the generalization capability of large language models.
📝 Abstract
Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision. We show that the weights favored by minimizing this loss need not yield better model performance on new tasks. In fact, we find that lower reconstruction loss can even degrade model performance on the same calibration data. Our analysis further shows that weights with lower reconstruction loss on calibration data can have higher loss than other weights when the distribution of input activations changes. Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions. DRQ refines the integer codes representing quantized weights within the existing quantization grid, keeping quantization parameters and inference operators unchanged. Extensive experiments show that DRQ improves models quantized by six representative PTQ methods, including AWQ, GPTQ, and ParoQuant, and delivers gains across both dense and mixture-of-experts large language models. These results establish DRQ as a general post-hoc refinement framework for weight-only PTQ, achieving better downstream performance without adding inference overhead.