OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

📅 2025-01-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the significant accuracy degradation in post-training quantization (PTQ) of large language models (LLMs), caused by skewed and heavy-tailed activation/weight distributions that widen optimal quantization ranges, this paper proposes a full-quantization method jointly optimizing weight and activation distribution alignment. Our contributions are threefold: (1) we introduce Quantization Space Utilization Ratio (QSUR), a novel metric to quantitatively evaluate distribution alignment; (2) we design a learnable orthogonal-plus-scaling equivalent transformation to enhance quantization robustness; and (3) we propose the KL-Top loss, which balances semantic fidelity and statistical robustness under limited calibration data. Under the W4A4KV4 configuration, our method reduces the performance gap with state-of-the-art methods by 32% on LLaMA-3-8B, while maintaining 99.5% of FP16 accuracy.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationNatural Language Processing: (Large) Language ModelsSearch and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Large language models for search
📝 Abstract
Post-training quantization (PTQ) has emerged as a widely adopted technique for compressing and accelerating Large Language Models (LLMs). The major challenge in LLM quantization is that uneven and heavy-tailed data distributions can expand the quantization range, thereby reducing bit precision for most values. Recent methods attempt to eliminate outliers and balance inter-channel differences by employing linear transformations; however, they remain heuristic and are often overlook optimizing the data distribution across the entire quantization space.In this paper, we introduce Quantization Space Utilization Rate (QSUR), a novel metric that effectively assesses the quantizability of transformed data by measuring the space utilization of the data in the quantization space. We complement QSUR with mathematical derivations that examine the effects and limitations of various transformations, guiding our development of Orthogonal and Scaling Transformation-based Quantization (OSTQuant). OSQuant employs a learnable equivalent transformation, consisting of an orthogonal transformation and a scaling transformation, to optimize the distributions of weights and activations across the entire quantization space. Futhermore, we propose the KL-Top loss function, designed to mitigate noise during optimization while retaining richer semantic information within the limited calibration data imposed by PTQ. OSTQuant outperforms existing work on various LLMs and benchmarks. In the W4-only setting, it retains 99.5% of the floating-point accuracy. In the more challenging W4A4KV4 configuration, OSTQuant reduces the performance gap by 32% on the LLaMA-3-8B model compared to state-of-the-art methods. href{https://github.com/BrotherHappy/OSTQuant}{https://github.com/BrotherHappy/OSTQuant}.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Post-Training Quantization
Data Distribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

OSTQuant
KL-Top loss function
Orthogonal and Scaling Transform
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xing Hu
Houmo AI
Y
Yuan Cheng
Houmo AI, Nanjing University
D
Dawei Yang
Houmo AI
Z
Zukang Xu
Houmo AI
Zhihang Yuan
Zhihang Yuan
Bytedance
Efficient AIModel CompressionMLLM
Jiangyong Yu
Jiangyong Yu
houmo.ai
C
Chen Xu
Houmo AI
Z
Zhe Jiang
Southeast University
Sifan Zhou
Sifan Zhou
Southeast University
RoboticsM/LLMsSpatial AIQuantization