🤖 AI Summary
This study addresses the challenge that outlier channels dominate quantization error in low-bit settings, where existing methods often induce weight bottlenecks or incur additional online overhead. We propose a training-free framework that introduces a novel channel permutation structure to transform scattered activations into contiguous memory accesses, combined with offline residual compensation to enable full precomputation. This design supports single-step GEMM inference under the MXFP4 format. By eliminating online computational bottlenecks, the proposed method achieves up to 2.6× operator-level speedup. Extensive evaluations on the Qwen3 series of large language models demonstrate that our approach preserves near-native decoding efficiency while surpassing existing baselines in accuracy.
📝 Abstract
Low-bit quantization of linear layers is often dominated by a small number of outlier channels. Existing smoothing, rotation, and residual-based methods can mitigate this issue, but may shift the quantization bottleneck to weights or introduce costly online operations. To address these limitations, we propose PRQuant (Permutation Residual Quantization), a training-free framework that combines channel permutation with offline weight residual compensation. PRQuant identifies the scaled-weight columns with the largest quantization errors and permutes them into contiguous tail blocks. This structure allows the corresponding weight residuals to be precomputed entirely offline, while replacing scattered activation gathering with simple contiguous access during inference, yielding a single regular MXFP4 GEMM for compensated computation. Experiments show that PRQuant substantially reduces down-projection reconstruction error, with scaling and residual compensation providing the main numerical gains while permutation enables a hardware-friendly contiguous layout. Comprehensive experiment results on Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507 illustrate that PRQuant achieves up to averagelly 2.6x and 1.8x operator speedup over BF16 respectively, while preserving near plain MXFP4 end-to-end decoding efficiency. Across five downstream benchmarks, PRQuant achieves the best average accuracy among the quantized methods, improving accuracy over MXFP4 by 1.24 and 0.55, respectively.