Tailoring the Quantization Space for 1-Bit KV Cache Compression

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the KV cache memory bottleneck in long-context LLM inference, where existing quantization methods suffer significant performance degradation under extreme 1-bit compression. To overcome this, we propose TaSQ, an efficient 1-bit KV cache compression method leveraging a tailored vector quantization target space. Its core innovations include query-guided weighting, cross-head normalization, and covariance-aware grouping to precisely model activation statistics, alongside RoPE-compatible transformations and adaptive channel grouping to preserve representational capacity under extreme compression. Implemented within the SGLang framework, TaSQ outperforms existing low-bit baselines across multiple benchmarks, enabling a 14× larger batch size and achieving a 1.87× higher peak throughput than BF16 while maintaining stable inference.
📝 Abstract
The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade substantially in the 1-bit regime. At such extreme compression, each codebook must represent a larger group of channels with a limited set of centroids, making effective use of its capacity increasingly challenging. To address this, we introduce $\textbf{TaSQ}$, which tailors the VQ target space by combining query-guided channel weighting, cross-head normalization, and covariance-aware channel grouping to better reflect the error sensitivity and statistical structure of cached activations. Since these transforms are RoPE-compatible and can be easily merged into projection weights and codebooks, TaSQ preserves the conventional VQ lookup structure and adds negligible serving overhead. Across general, long-chain-of-thought reasoning, and long-context retrieval benchmarks, TaSQ consistently outperforms existing low-bit KV cache VQ baselines while preserving reasoning stability. On a single RTX 6000 Ada GPU, its SGLang implementation supports up to $14\times$ larger batch sizes and achieves $1.87\times$ higher peak throughput compared to the BF16 baseline.
Problem

Research questions and friction points this paper is trying to address.

KV cache compression
vector quantization
1-bit quantization
long-context LLM inference
memory bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vector Quantization
1-Bit KV Cache Compression
Query-Guided Channel Weighting
Covariance-Aware Channel Grouping
RoPE-Compatible
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Minsoo Cheong
Seoul National University
D
Donghyun Son
Stanford University
Sungjoo Yoo
Sungjoo Yoo
Seoul National University
memorystorage