🤖 AI Summary
To address the significant accuracy degradation in post-training quantization (PTQ) of large language models (LLMs), caused by skewed and heavy-tailed activation/weight distributions that widen optimal quantization ranges, this paper proposes a full-quantization method jointly optimizing weight and activation distribution alignment. Our contributions are threefold: (1) we introduce Quantization Space Utilization Ratio (QSUR), a novel metric to quantitatively evaluate distribution alignment; (2) we design a learnable orthogonal-plus-scaling equivalent transformation to enhance quantization robustness; and (3) we propose the KL-Top loss, which balances semantic fidelity and statistical robustness under limited calibration data. Under the W4A4KV4 configuration, our method reduces the performance gap with state-of-the-art methods by 32% on LLaMA-3-8B, while maintaining 99.5% of FP16 accuracy.
📝 Abstract
Post-training quantization (PTQ) has emerged as a widely adopted technique for compressing and accelerating Large Language Models (LLMs). The major challenge in LLM quantization is that uneven and heavy-tailed data distributions can expand the quantization range, thereby reducing bit precision for most values. Recent methods attempt to eliminate outliers and balance inter-channel differences by employing linear transformations; however, they remain heuristic and are often overlook optimizing the data distribution across the entire quantization space.In this paper, we introduce Quantization Space Utilization Rate (QSUR), a novel metric that effectively assesses the quantizability of transformed data by measuring the space utilization of the data in the quantization space. We complement QSUR with mathematical derivations that examine the effects and limitations of various transformations, guiding our development of Orthogonal and Scaling Transformation-based Quantization (OSTQuant). OSQuant employs a learnable equivalent transformation, consisting of an orthogonal transformation and a scaling transformation, to optimize the distributions of weights and activations across the entire quantization space. Futhermore, we propose the KL-Top loss function, designed to mitigate noise during optimization while retaining richer semantic information within the limited calibration data imposed by PTQ. OSTQuant outperforms existing work on various LLMs and benchmarks. In the W4-only setting, it retains 99.5% of the floating-point accuracy. In the more challenging W4A4KV4 configuration, OSTQuant reduces the performance gap by 32% on the LLaMA-3-8B model compared to state-of-the-art methods. href{https://github.com/BrotherHappy/OSTQuant}{https://github.com/BrotherHappy/OSTQuant}.