Align, Then Correct: Training-Free Two-Stage Low-Rank Compensation for Extremely Quantized Large Language Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the accuracy degradation caused by symmetric calibration and high-rank compensation in ultra-low-bit quantization. We propose a training-free, two-stage low-rank compensation framework that introduces a Fisher-weighted asymmetric objective. The method first aligns layer outputs to compress rank requirements, then absorbs first-order error signals via natural gradient descent, overcoming the limitations of conventional second-order approximations while yielding a closed-form solution. By integrating truncated SVD with the QuIP# quantizer, our approach significantly reduces perplexity on Qwen models, recovering 51%–84% of the performance gap relative to FP16 precision and comprehensively outperforming existing baseline methods.
📝 Abstract
Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without any training. We show that existing compensators are limited by two shared simplifications. They calibrate symmetrically, evaluating the full-precision and compensated weights on the same activation, which yields a compensation target that is inherently high-rank -- so a fixed rank budget captures only a small fraction of it. And they minimize only the second-order term of the loss, although the compensated model is not stationary: a first-order descent direction larger than the applied compensation itself remains in every layer, and no reconstruction objective can absorb it. We propose a two-stage closed-form framework that removes both simplifications. Stage 1 aligns each layer's output with the full-precision model under a Fisher-weighted asymmetric objective, concentrating the rank budget on a rank-compressible target. Stage 2 re-measures statistics on the compensated model and applies a rank-constrained natural-gradient step that absorbs the remaining first-order signal. Every adapter is the result of a single truncated SVD; backward passes serve only to collect statistics. At 2 bits under QuIP#, our method reduces WikiText-2 perplexity from 12.43 to 10.26 on Qwen3-8B and from 21.11 to 13.22 on Qwen3-4B. On the held-out C4 corpus, it recovers 51% and 84% of the gap to FP16, versus 31% and 63% for the strongest baseline, with consistent gains in the seven-task zero-shot average, at higher bit-widths, and under a distinct quantizer.
Problem

Research questions and friction points this paper is trying to address.

extreme quantization
low-rank error compensation
large language models
asymmetric calibration
first-order residual
Innovation

Methods, ideas, or system contributions that make the work stand out.

Low-rank quantization error compensation
Training-free adaptation
Two-stage closed-form framework
Asymmetric Fisher-weighted alignment
Rank-constrained natural gradient
💼 Related Jobs
No related jobs found.
S
Seobin Song
Hanyang University, Seoul, Republic of Korea
G
Geonho Lee
Hanyang University, Seoul, Republic of Korea
J
Janghwan Lee
Qualcomm AI Research, Qualcomm Korea YH, Seoul, Republic of Korea
Jungwook Choi
Jungwook Choi
Hanyang University
Deep Neural NetworkQuantizationLarge Language ModelEfficient AIAI Accelerator