DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation

📅 2026-07-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the substantial memory overhead of Key-Value (KV) caching in long-context inference with large language models. It proposes the first low-rank compression framework that treats Keys and Values differently: for the Key cache, it dynamically groups attention heads based on Centered Kernel Alignment (CKA) and adaptively allocates rank budgets; for the Value cache, it employs offline calibration to optimize low-rank decomposition. Evaluated on three instruction-tuned large language models, the method achieves significant compression of Key cache parameters while maintaining competitive accuracy, demonstrating particularly strong performance in multi-head attention architectures.
📝 Abstract
As the inference phase of Large Language Models (LLMs) requires handling long context windows, the Key-Value (KV) cache initially appears to address this challenge but eventually becomes a significant bottleneck as the context window continues to grow. Low-rank compression has recently been studied as an effective approach to reduce KV cache memory while maintaining model performance. However, only a few existing methods treat the Key and Value caches differently, despite their distinct roles. Moreover, these methods typically employ fixed attention-head grouping, which may not fully exploit the structural similarity among attention heads. In this paper, we propose an improved low-rank KV cache compression framework. For the Key cache, we dynamically group attention heads based on Centered Kernel Alignment (CKA) similarity and allocate the rank budget adaptively under a parameter budget. For the Value cache, we adopt the same approach as ReCalKV, refining the low-rank decomposition through offline calibration to improve reconstruction quality. Experimental results on three instruction-tuned LLMs show that our method reduces the number of Key cache parameters while maintaining competitive accuracy. We further observe that the proposed strategy is particularly effective for Multi-Head Attention (MHA) models, whereas it should be applied more conservatively to Grouped-Query Attention (GQA) models, especially in long-context settings.
Problem

Research questions and friction points this paper is trying to address.

KV cache compression
attention head grouping
adaptive rank allocation
large language models
low-rank decomposition
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV cache compression
dynamic head grouping
adaptive rank allocation
low-rank decomposition
Centered Kernel Alignment
T
Tan T. Nguyen
Full Stack Data Science
Q
Quan V. Dang
Department of Computer Science, University College London