G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation

πŸ“… 2026-09-25
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses critical limitations in post-training quantization for large language models, including the lack of global supervision in local optimization, the neglect of first-order gradients, and stale Hessian approximations. To overcome these challenges, we propose a unified quantization framework based on generalized gradient compensation. Methodologically, block-wise optimization is employed to dynamically refresh gradient and Hessian approximations, thereby eliminating information staleness. Furthermore, a trust-region scaling mechanism is introduced to stabilize first-order compensation and prevent divergence during weight updates. Extensive experiments demonstrate that the proposed framework significantly outperforms existing state-of-the-art methods across diverse model architectures and bit-width configurations, enabling quantized models to more closely align with full-precision baselines in performance.
πŸ“ Abstract
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds. This paper presents G$^2$PTQ, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective. By refreshing gradient and Hessian estimates before quantizing each Transformer block, G$^2$PTQ avoids the staleness of prior global methods. Furthermore, to stabilize the exact first-order compensation, we introduce a trust-region scaling mechanism that dynamically bounds the gradient step to prevent exploding weight updates. Finally, we derive efficient implementations for block-wise Hessian approximation and exact gradient compensation. Experimental results on various model families and bit-widths demonstrate that G$^2$PTQ enables better alignment with the full-precision model, outperforming state-of-the-art baselines. Code is available at: https://github.com/G2PTQ/G2PTQ.
Problem

Research questions and friction points this paper is trying to address.

Post-training quantization
Large language models
Hessian estimation
Gradient compensation
Global supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-Training Quantization
Generalized Gradient Compensation
Trust-Region Scaling
Block-wise Optimization
Large Language Models
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
R
Ruikang Liu
ZTE Corporation
Haoli Bai
Haoli Bai
Huawei Technologies
natural language processingmodel compression
Y
Yuxuan Sun
Northwestern Polytechnical University
Q
Qian Zhang
Peking University
W
Wenzheng Cai
ZTE Corporation
Y
Yanqi Hao
ZTE Corporation
Feiyu Wang
Feiyu Wang
Fudan University
computer vision
W
Weidong Zhong
ZTE Corporation
Z
Zhuang Wang
ZTE Corporation
Tong Yang
Tong Yang
Peking University, Beijing, China. PKU. εŒ—δΊ¬ε€§ε­¦
SketchNetwork measurementBloom filterIP lookupHash Table
X
Xiangsheng Zhou
ZTE Corporation, Nanjing University of Aeronautics and Astronautics