Page-EntroKV: Hardware-Aligned, Entropy-Weighted KV-Cache Eviction under Grouped-Query Attention

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses memory bloat and retrieval-head dilution in KV cache eviction for Grouped Query Attention (GQA) architectures, which arise from independent head-wise selection. We propose a hardware-aligned, entropy-weighted eviction framework that aggregates scoring at the physical group granularity and projects it onto page frames for execution. By leveraging Rényi-2 entropy to isolate sink heads, we establish a theoretical proof of the joint overhead ratio, ensuring strict budget preservation and precise page accounting. Experimental results demonstrate that our framework eliminates 4.75× redundancy and reduces sink masquerading by 13×, achieving 100% needle recall under a 20% budget constraint and significantly outperforming mean-pooling baselines.
📝 Abstract
Serving long-context autoregressive language models is constrained by the key-value (KV) cache. Most dynamic eviction methods score token importance per query head and choose tokens independently. This fits poorly with grouped-query attention (GQA), where several query heads share one physical KV buffer: divergent per-head selections force the serving engine to retain the union of their choices - inflating the cache by up to the group ratio r - while arithmetic mean pooling dilutes the specialized retrieval heads that carry factual recall. We introduce Page-EntroKV, a formal framework for KV-cache eviction operating at the granularity GQA serving actually allocates. Heads within each physical group are pooled by weights derived from sink-isolated collision (Renyi-2) entropy - one inner product per head, computed once at prefill with no calibration - so sink heads cannot masquerade as retrieval heads. Pooled scores are projected onto PagedAttention page frames, and eviction executes at the hardware tuple (layer, group, page). We formalize the union overhead ratio (UOR) and intra-group disagreement, prove an exact identity linking them for two-head groups alongside two-sided bounds at every group ratio, prove strict budget preservation and a finite-context needle-retention bound that arithmetic mean pooling provably violates, and give exact per-layer page accounting. On a pilot architecture (Qwen2.5-1.5B-Instruct, r=6), head-independent replay over 2,240 group measurements yields union overhead up to 4.75x at a 2% budget, while Page-EntroKV holds UOR exactly 1.000; sink isolation removes a 13x sink masquerade; needle recall is 100% versus 0% for mean pooling at a 20% budget; retained cardinality is exact for every page size; and QA and code tasks remain solvable at 20% retention.
Problem

Research questions and friction points this paper is trying to address.

KV-cache eviction
Grouped-Query Attention
long-context language models
union overhead
sink token
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV-Cache Eviction
Grouped-Query Attention
Renyi-2 Entropy
PagedAttention
Union Overhead Ratio
I
Inbasekaran S
Department of Computer Science and Engineering, SRM Institute of Science and Technology, India