WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the memory and bandwidth bottlenecks of KV caches in long-context reasoning by proposing WUSH-KV, a low-bit quantization scheme. It leverages second-order statistics to construct data-adaptive transformations that minimize quantization error. Notably, this work innovatively decouples key and value transformations, folding the value transformation into the weight matrix, and proves its near-optimality under specific quantizers. The method integrates RoPE post-processing and percentile-clipped affine quantization within the SGLang engine. Experimental results demonstrate that WUSH-KV substantially reduces inter-layer reconstruction error and achieves the lowest perplexity at 2-bit precision, outperforming the OSCAR baseline.
πŸ“ Abstract
KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one such quantizer, QuEST INT, we show that, under mild assumptions, the WUSH transform is near-optimal. With this quantizer, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among other tested transforms. For end-to-end evaluation, we integrate WUSH-KV into SGLang using OSCAR-style percentile-clipped affine quantization. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks.
Problem

Research questions and friction points this paper is trying to address.

KV cache quantization
long-context inference
memory bottleneck
low-bit quantization
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV Cache Quantization
Data-Adaptive Transforms
Low-bit Quantization
Long-context Inference
Second-order Statistics
πŸ”Ž Similar Papers