PRQuant: Permutation Residual Quantization for Low-Overhead Inference

📅 2026-08-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that outlier channels dominate quantization error in low-bit settings, where existing methods often induce weight bottlenecks or incur additional online overhead. We propose a training-free framework that introduces a novel channel permutation structure to transform scattered activations into contiguous memory accesses, combined with offline residual compensation to enable full precomputation. This design supports single-step GEMM inference under the MXFP4 format. By eliminating online computational bottlenecks, the proposed method achieves up to 2.6× operator-level speedup. Extensive evaluations on the Qwen3 series of large language models demonstrate that our approach preserves near-native decoding efficiency while surpassing existing baselines in accuracy.
📝 Abstract
Low-bit quantization of linear layers is often dominated by a small number of outlier channels. Existing smoothing, rotation, and residual-based methods can mitigate this issue, but may shift the quantization bottleneck to weights or introduce costly online operations. To address these limitations, we propose PRQuant (Permutation Residual Quantization), a training-free framework that combines channel permutation with offline weight residual compensation. PRQuant identifies the scaled-weight columns with the largest quantization errors and permutes them into contiguous tail blocks. This structure allows the corresponding weight residuals to be precomputed entirely offline, while replacing scattered activation gathering with simple contiguous access during inference, yielding a single regular MXFP4 GEMM for compensated computation. Experiments show that PRQuant substantially reduces down-projection reconstruction error, with scaling and residual compensation providing the main numerical gains while permutation enables a hardware-friendly contiguous layout. Comprehensive experiment results on Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507 illustrate that PRQuant achieves up to averagelly 2.6x and 1.8x operator speedup over BF16 respectively, while preserving near plain MXFP4 end-to-end decoding efficiency. Across five downstream benchmarks, PRQuant achieves the best average accuracy among the quantized methods, improving accuracy over MXFP4 by 1.24 and 0.55, respectively.
Problem

Research questions and friction points this paper is trying to address.

Low-bit quantization
Outlier channels
Quantization bottleneck
Inference overhead
Large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Permutation Residual Quantization
Training-free Quantization
Offline Weight Compensation
Channel Permutation
MXFP4
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Peiran Wang
Huawei Technologies Co., Ltd.
A
Anqi Wang
Tsinghua University
Jiaying Zhao
Jiaying Zhao
Huawei Technologies Co., Ltd.
H
Huiwen Yang
Huawei Technologies Co., Ltd.
Z
Zhenyu Ming
Huawei Technologies Co., Ltd.
Y
Yuantian Shao
Nanjing University of Science and Technology
R
Rongqian Wang
Huawei Technologies Co., Ltd.
Yiwu Yao
Yiwu Yao
Peking University
Artificial Intelligence
Kun Tian
Kun Tian
Intel
X
Xin Yao
Huawei Technologies Co., Ltd.
G
Gong Zhang
Huawei Technologies Co., Ltd.
F
Fan Yang
Tsinghua University
Zhongyi Huang
Zhongyi Huang
Professor of mathematics, Tsinghua University
Scientific Computingmultiscale methodssingular perturbation problemshigh frequency waves