manage kv caches

Designs, implements, and evaluates systems that store, manage, and manipulate key–value activation caches for autoregressive attention, including compression, summarization, pruning, eviction, block‑granularity placement, and transfer scheduling to reduce memory, compute, and network cost while preserving causal masking semantics. Builds prefetching and batching strategies, lightweight calibration and test‑time correction mappings, and measurement pipelines to minimize end‑to‑end resume latency and control error accumulation from cached features.

managekvcaches

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.79
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$241K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the high memory overhead and low throughput in autoregressive image generation caused by caching all historical visual tokens, a problem exacerbated by existing methods that allocate a uniform KV cache budget across all attention heads despite their heterogeneous behaviors. To overcome this limitation, the paper proposes HeadKV, a head-aware KV cache compression framework that dynamically allocates cache budgets based on whether each attention head exhibits local or global contextual preferences. By identifying head types early in the sequence and applying a hierarchical token eviction strategy, HeadKV achieves efficient compression while preserving long-range dependencies. Notably, the method generalizes across inputs without requiring additional training or dataset-specific statistics. Experiments demonstrate that HeadKV substantially reduces memory consumption and boosts inference throughput across multiple autoregressive image generation models, all while maintaining high generation quality.

attention headsautoregressive image generationKV cache compression

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

May 26, 2025
KL
Kunjun Li
🏛️ National University of Singapore | University of Washington

Visual autoregressive (VAR) models suffer from exponential growth of key-value (KV) cache memory consumption and computational redundancy during multi-scale inference, due to their coarse-to-fine generation paradigm. To address this, we propose ScaleKV, a novel KV compression framework that introduces scale-aware layer functional partitioning—distinguishing between *drafters* (coarse-scale layers) and *refiners* (fine-scale layers)—enabling differentiated KV cache management across scales. ScaleKV integrates architectural analysis of Transformers, multi-scale attention modeling, dynamic KV pruning, and hierarchical cache capacity allocation. Evaluated on the Infinity text-to-image VAR model, ScaleKV reduces KV cache memory usage to 10% of the baseline while preserving pixel-level reconstruction fidelity. This work constitutes the first systematic solution to KV cache explosion in multi-scale VAR inference, establishing a new paradigm for efficient visual generation.

Coarse-to-fine VAR methods cause computational redundancy during inferenceExponential KV cache growth in VAR models increases memory useScaleKV optimizes cache by analyzing layer-specific scale attention patterns

This work addresses the substantial GPU memory consumption caused by the linear growth of key-value (KV) cache with context length in large language model inference, which severely limits long-context generation efficiency. The authors propose a training-free KV cache compression method that leverages anti-causal attention to compute a “surprise score” for each token, dynamically pruning redundant entries that are highly predictable from subsequent tokens. To further reduce computational overhead, they introduce a single-layer Transformer approximation applied exclusively to the final layer. By integrating anti-causal masking with online cache reuse, the approach achieves state-of-the-art or comparable performance across multiple open-source large language models and benchmarks, significantly reducing memory usage and accelerating inference while preserving generation quality.

cache managementcontext lengthKV cache

This work addresses the performance degradation in dynamic sparse attention mechanisms, where per-token selection of key-value cache entries leads to fragmented working sets and poor cache locality, thereby increasing last-level cache (LLC) misses and reducing decoding throughput. To mitigate this issue, the authors propose a service-oriented cache architecture featuring a lightweight indexer that analyzes key-value access patterns, coupled with an LLC reservation mechanism and a fine-grained, token-level LRU replacement policy. This design effectively alleviates cache fragmentation and improves data locality. Experimental evaluation across multiple open-source backbone models demonstrates that the proposed approach significantly reduces LLC stall misses and enhances online inference throughput for dynamic sparse attention models.

Cache LocalityDecode ThroughputDynamic Sparse Attention

This work addresses critical limitations in KV cache management for large-scale GPU inference—namely, fixed sizing, single-level storage, and passive eviction policies—that lead to excessive memory waste and redundant computation. The authors propose a unified KV cache management system that enables precise sizing for diverse attention architectures, including MLA, and introduces a six-tier heterogeneous memory hierarchy spanning from HBM to parallel file systems. A key innovation is a block-type-aware Bayesian reuse prediction model based on Beta conjugate priors, which facilitates head-granularity eviction and RoPE-aware prefetching. Experimental results demonstrate that the system achieves 70–84% cache hit rates, reduces first-token latency by 1.4–2.1×, improves throughput by 1.7–2.9×, lowers inference costs by 47%, and delivers an effective cache capacity of 38 TB per node.

attention architectureKV cachelarge-scale GPU inference

Latest Papers

What's happening recently
View more

This work addresses the substantial memory and bandwidth bottlenecks in autoregressive decoding caused by the linear growth of key-value (KV) cache with context length. Existing KV cache eviction methods rely on static heuristics or proxy scores that inadequately estimate each cache entry’s contribution to future token generation, often resulting in significant performance degradation. To overcome this limitation, the authors propose a supervised learning framework that directly optimizes KV cache eviction under a fixed budget by leveraging future attention targets as supervision signals. They further introduce a delayed memory scorer that implicitly guides online pruning using near-future context, eliminating the need for explicit computation of dense attention maps. Evaluated on Qwen3-4B and Qwen3-8B, the method retains 97%–98% of original model performance at aggressive compression rates of 75%–88%, substantially outperforming current baselines.

autoregressive decodingcache evictionfuture token utility

This study addresses memory bloat and retrieval-head dilution in KV cache eviction for Grouped Query Attention (GQA) architectures, which arise from independent head-wise selection. We propose a hardware-aligned, entropy-weighted eviction framework that aggregates scoring at the physical group granularity and projects it onto page frames for execution. By leveraging Rényi-2 entropy to isolate sink heads, we establish a theoretical proof of the joint overhead ratio, ensuring strict budget preservation and precise page accounting. Experimental results demonstrate that our framework eliminates 4.75× redundancy and reduces sink masquerading by 13×, achieving 100% needle recall under a 20% budget constraint and significantly outperforming mean-pooling baselines.

Grouped-Query AttentionKV-cache evictionlong-context language models

This study addresses the trade-off between KV cache quantization and linear attention in Transformers regarding storage and computational costs, highlighting the absence of a unifying mechanism. To bridge this gap, we propose RAM-Net, which unifies per-KV compression and multi-KV aggregation through soft allocation over a discrete address space. Theoretically, we demonstrate that soft address allocation generalizes hard quantization matching and supports recurrent state updates, thereby establishing a principled weight transfer pathway from pretrained Transformers to RAM-Net. Empirically, across nine pretrained models, RAM-Net recovers 87.1% of the teacher model’s average accuracy gains with a fine-tuning budget of merely 500M tokens, achieving highly efficient attention.

efficient attentionKV-cache quantizationlinear attention

This work addresses the high computational and storage costs, methodological fragmentation, and lack of a unified theoretical foundation in memory management for large language model agents. Framing multi-level memory compression through rate-distortion theory, the study formulates it as an optimization problem of preserving task-critical information under resource constraints. It introduces a general compression objective and a layer-agnostic lower bound, establishes a seven-dimensional taxonomy, and achieves, for the first time, cross-layer transfer of compression mechanisms. The analysis reveals fundamental limitations of attention magnitude and temporal decay as universal retention signals. Furthermore, the authors develop a unified evaluation benchmark with reference experiments, distill design principles for memory compression in multi-turn interactions, and systematically outline open challenges in the field.

agent memorylarge language modelsmemory compaction

Hot Scholars

EX

Enze Xie

NVIDIA Research, MMLab@HKU
computer visiongenerative AI
SK

Souvik Kundu

Sr. Staff Research Scientist, Intel AI Group; Ph.D - USC; IEEE/ACM DAC under-40 Innovator
Efficient AIEnergy Efficient ComputingLLMMultimodal Foundation Models
XH

Xuming Hu

Assistant Professor, HKUST(GZ) / HKUST
Natural Language ProcessingLarge Language Model
KD

Kuntai Du

University of Chicago
Large Language ModelsVideo analytics