Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the substantial storage and access overhead of long-context KV caches, noting that existing rotation-based quantization methods reduce bit-width but disperse query energy, thereby hindering efficient channel pruning. To overcome this limitation, this work proposes the Dual-QK framework, which leverages paired non-orthogonal transformations to balance key quantization scales and concentrate query energy. By integrating calibration statistics, partial key whitening, and a Channel-0 protection mechanism, Dual-QK achieves synergistic optimization of low-bit quantization and dynamic channel pruning. Experimental results demonstrate that the proposed method outperforms OSCAR in accuracy at 40% sparsity. Furthermore, on 128K context lengths, it attains 6.8× compression and an 8.3× reduction in memory reads, yielding up to a 3.75× improvement in decoding throughput.
📝 Abstract
Long inputs and extended generation increase the storage and access costs of the key-value (KV) cache. Low-bit quantization reduces storage and memory traffic, while query-channel pruning can further reduce key-cache reads. Rotation-based quantization redistributes the energy of key outliers across channels. To maintain computational invariance, the same orthogonal transform must be applied to queries, preserving query-key dot products. However, this rotation can disperse query energy, weakening the separation between a few large components to retain and many small ones to prune. We introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict. Using calibrated query and key statistics, Dual-QK combines partial key whitening with a query-aligned basis to balance key scales for INT2 quantization and concentrate query energy for dynamic channel pruning. Channel-0 protection and bucket-relative RoPE support low-bit accuracy over long contexts. Experiments on four models across five generative benchmarks and long-context retrieval tasks show improved accuracy over OSCAR on most tasks at 40% query-channel sparsity. At a 128K context, Dual-QK provides $6.8\times$ KV-cache compression and an estimated $8.3\times$ reduction in KV read volume relative to unpruned BF16. Under the evaluated configurations, our SGLang implementation achieves up to $3.75\times$ the decoding throughput of unpruned BF16.
Problem

Research questions and friction points this paper is trying to address.

KV cache compression
low-bit quantization
query-channel pruning
energy dispersion
long-context generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Non-orthogonal transforms
INT2 KV cache quantization
Dynamic channel pruning
Partial key whitening
Bucket-relative RoPE
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sunjoo Whang
KAIST
J
Jungjun Oh
KAIST
M
Minsung Kim
KAIST
D
Dongho Seo
GIST
Jisu Shin
Jisu Shin
GIST AI Graduate School
Computer Vision
G
Gregory Kielian
Google Research
Hoi-Jun Yoo
Hoi-Jun Yoo
Professor of Electrical Engineering, KAIST
S
Sangjin Kim
GIST