ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory bottleneck in recurrent Transformers caused by the linear growth of KV cache with recurrence steps, as well as the severe accuracy degradation associated with existing low-bit quantization methods. To overcome these limitations, this work proposes the first residual quantization paradigm tailored for recurrent architectures. By exploiting inter-recurrence KV state similarity, the method stores low-bit residuals using the final state as a reference, achieving efficient INT2 quantization through least-squares scaling, orthogonal rotation, and layer-wise mixed precision. Experimental results demonstrate that the proposed approach reduces KV memory consumption by 80.7% and improves accuracy by 13% under equivalent memory budgets, while accelerating decoding speed and peak throughput by 2.73× and 4.15×, respectively.
📝 Abstract
Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count. However, KV cache memory still scales with the number of loops, becoming a key memory bottleneck that limits batch size and inference throughput. KV cache quantization can alleviate this bottleneck, but existing methods often suffer substantial accuracy degradation at aggressive low-precision regimes. We observe that looped Transformers offer a unique opportunity: KV states across loops are highly similar. Based on this observation, we propose ResidualQuant, which uses the final-loop KV states as a reference and represents the remaining loops with low-precision residuals. Our method further combines least-square scaling and rotations applied to the residuals, as well as loop-wise mixed precision, to enable accurate quantization down to INT2 while retaining efficient reconstruction. Across multiple looped Transformer models and mathematical reasoning and code generation benchmarks, ResidualQuant consistently improves the accuracy-memory tradeoff over state-of-the-art rotation-based KV quantization. In particular, our method retains accuracy close to BF16 under mixed-precision settings while reducing theoretical KV storage by 80.7%, achieving up to 13.0% higher accuracy than the rotation-based baseline at the same memory budget. On an RTX 5090, the reduced KV memory traffic improves fixed-batch decode throughput by up to 2.73x, while the smaller memory footprint enables up to 2x larger batches, improving peak throughput by up to 4.15x.
Problem

Research questions and friction points this paper is trying to address.

Looped Transformers
KV Cache Quantization
Memory Bottleneck
Low-Precision Quantization
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV Cache Quantization
Looped Transformers
Residual Quantization
Mixed Precision
INT2
🔎 Similar Papers
2024-01-31Neural Information Processing SystemsCitations: 193