page-aware kv decoding

Design and implement decoder runtime components that execute autoregressive decode workflows while managing key-value (KV) cache and attention with page-aware memory management and scheduling. Build adaptive kernel selection, block-table decode-attention implementations, persistent KV tiling, and workqueue scheduling that maps work by KV-head groups to minimize KV-cache movement and to support native page-table layouts.

page-awarekvdecoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

A Survey on Large Language Model Acceleration based on KV Cache Management

Dec 27, 2024
HL
Haoyang Li
🏛️ The Hong Kong Polytechnic University | Hong Kong University of Science and Technology | Huazhong University of Science and Technology | The Chinese University of Hong Kong | Nanyang Technological University

Large language models (LLMs) suffer from high memory overhead and inefficiency in long-context and real-time inference scenarios due to KV cache accumulation. To address this, we propose the first holistic, three-tiered KV cache management framework—operating at the token, model, and system levels—that unifies cache selection, quantization, low-rank decomposition, attention sparsification/windowing, and hardware-aware scheduling. We establish a standardized benchmark covering both text and multimodal tasks, release the first open-source repository for KV cache management research (Awesome-KV-Cache-Management), and provide a comprehensive technical taxonomy with empirical comparisons across methods. This work advances the systematization and standardization of KV cache management methodologies, significantly improving inference efficiency and deployment feasibility of LLMs under resource constraints.

Computational DemandLarge Language ModelsMemory Requirement

This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.

cache managementdistributed memoryKV cache

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the performance bottleneck caused by KV cache movement during long-context large language model (LLM) inference, particularly under mixed sequence lengths and low token activity, where existing scheduling strategies struggle to efficiently utilize consumer-grade GPU resources. The paper proposes PersistentKV, a decoding scheduler for paged KV caches that integrates page-aware scheduling with adaptive mechanisms for the first time. By mapping tasks according to KV head groups, reusing K/V blocks, leveraging native page table support, and maintaining compact work queues, PersistentKV enables fine-grained, non-empty-task-driven scheduling. Experiments on an RTX 3060 demonstrate throughput improvements of 1.06–1.40× over FlashInfer, with no performance degradation in boundary cases, underscoring the critical impact of task assignment on LLM serving efficiency.

commodity GPUsdecode schedulingKV cache

This work addresses the high memory consumption of key-value (KV) caching in visual autoregressive (VAR) models for image generation, which often exceeds several gigabytes and limits practical deployment. The authors propose a fine-grained, attention-head-specific KV cache compression method that leverages offline calibration to assess each head’s reliance on historical tokens. By integrating attention-score-based head ranking, static pruning, and dynamic cache allocation, the approach enables differentiated compression under a fixed memory budget. Evaluated on the Infinity-2B model, this technique achieves a 2× higher KV cache compression ratio compared to existing methods while preserving or even improving image fidelity, prompt alignment, and human-perceived quality, thereby establishing a new state of the art in VAR model cache compression.

attention headsimage generationKV-cache compression

Existing KV cache compression strategies for long-context large language models are rigid and neglect task-specific and layer-wise heterogeneity. Method: This paper proposes a dynamic hierarchical adaptive compression mechanism featuring (1) task-aware dynamic budget allocation, enabling per-layer, online adjustment of retained token counts; and (2) lightweight cache reconfiguration guided by inter- and intra-layer activation pattern analysis, under global and layer-specific budget constraints. The method requires no fine-tuning and performs periodic optimization during inference. Contribution/Results: With only 1.7% of the original KV cache retained, our approach achieves 85% of full-cache LongBench performance; under extreme compression (0.9% cache), it surpasses state-of-the-art methods by 11% accuracy on the Needle-in-a-Haystack task. The method significantly improves the trade-off between inference efficiency and accuracy for long-context processing.

Adaptive KV cache compressionEfficient long-context LLM performanceTask-specific token retention optimization

Existing lossless KV cache management approaches overlook the computational efficiency of GPU attention kernels, resulting in high inference latency. This work proposes AsymCache, a system that, for the first time, incorporates GPU computation latency into cache eviction decisions. By integrating multi-segment attention mechanisms, a position-aware recomputation cost model, and adaptive chunked scheduling, AsymCache jointly optimizes cache hit rates and computational efficiency. Compared to state-of-the-art baselines, AsymCache reduces first-token latency by 1.90–2.03× and per-token generation time by 1.62–1.71×, while achieving an average 18.1% reduction in job latency when deployed in the Continuum system.

computation latencyGPU attention kernel efficiencyKV-cache management

SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation

Dec 18, 2024
JW
Jialong Wu
🏛️ Southeast University | King's College London | The Alan Turing Institute

To address memory and bandwidth bottlenecks induced by KV caching in long-context generation, this paper proposes the first dual-stage separation compression paradigm—distinctly optimizing prefilling and decoding. During prefilling, excessive compression is avoided to preserve contextual understanding; during decoding, a sliding-window re-hitter identification mechanism coupled with adaptive discontinuous memory transfer dynamically retains critical key-value pairs. The method is plug-and-play compatible with mainstream KV compression techniques. Evaluated on LongGenBench, it achieves significant reductions in memory footprint and memory bandwidth consumption compared to prefilling-only compression baselines, while maintaining lossless generation quality and strong generalization across diverse long-context tasks.

Addressing decoding phase neglect in KV cache optimizationOptimizing KV cache compression for long-context generationReducing memory usage in long-output generation tasks

Latest Papers

What's happening recently
View more

This study reveals a pronounced non-monotonic latency behavior in Apple’s Metal Performance Shaders (MPS) backend during autoregressive decoding, challenging the widely held assumption that KV caching universally improves inference efficiency. Through controlled experiments on models including GPT-2, BLOOM, and OPT, the authors systematically compare MPS against CPU and NVIDIA T4 platforms, uncovering latency spikes of up to 21× within specific decoding length ranges on MPS—phenomena unexplained by memory pressure or prefill costs. The work provides the first characterization of the discontinuous latency patterns in MPS decoding, attributing them to backend execution scheduling mechanisms. These findings underscore the critical importance of hardware-aware evaluation for inference optimization and caution that aggregate benchmarking metrics may obscure such pivotal performance anomalies.

Apple MPSautoregressive decodinginference performance

This work addresses memory inefficiency and tail latency in static-graph large language model serving, caused by heterogeneous request lengths, asynchronous completion, and fragmented KV cache layouts. The authors propose KV-RM, a runtime system that decouples logical history from physical storage to standardize KV cache movement under fixed decoding interfaces. By employing block-level paging to track active states and designing a coalesced transfer path that aggregates non-contiguous KV mappings into large contiguous blocks, KV-RM aligns with fixed-shape attention kernels. This approach requires neither distant-history summarization nor dynamic scheduling, significantly improving throughput and tail latency under mixed-length request loads, reducing reserved memory overhead, and eliminating bursty latency spikes during production replay.

KV-cachememory fragmentationonline decoding

This work addresses the substantial memory and bandwidth bottlenecks in autoregressive decoding caused by the linear growth of key-value (KV) cache with context length. Existing KV cache eviction methods rely on static heuristics or proxy scores that inadequately estimate each cache entry’s contribution to future token generation, often resulting in significant performance degradation. To overcome this limitation, the authors propose a supervised learning framework that directly optimizes KV cache eviction under a fixed budget by leveraging future attention targets as supervision signals. They further introduce a delayed memory scorer that implicitly guides online pruning using near-future context, eliminating the need for explicit computation of dense attention maps. Evaluated on Qwen3-4B and Qwen3-8B, the method retains 97%–98% of original model performance at aggressive compression rates of 75%–88%, substantially outperforming current baselines.

autoregressive decodingcache evictionfuture token utility

This work addresses the substantial memory overhead of Key-Value (KV) caching in long-context inference with large language models. It proposes the first low-rank compression framework that treats Keys and Values differently: for the Key cache, it dynamically groups attention heads based on Centered Kernel Alignment (CKA) and adaptively allocates rank budgets; for the Value cache, it employs offline calibration to optimize low-rank decomposition. Evaluated on three instruction-tuned large language models, the method achieves significant compression of Key cache parameters while maintaining competitive accuracy, demonstrating particularly strong performance in multi-head attention architectures.

adaptive rank allocationattention head groupingKV cache compression

Hot Scholars

AM

Alexander Moreno

Institute of Foundation Models, MBZUAI
LLM pre-trainingtraining dynamicsfoundation models
SH

Shibo Hao

Ph.D. student, UC San Diego
machine learninglarge language model
TW

Taylor W. Killian

Senior Research Scientist, MBZUAI Institute of Foundation Models
Machine LearningReinforcement LearningHealthcareTransfer Learning
ZF

Zhenman Fang

Associate Professor, Simon Fraser University
Hardware AccelerationFPGAsReconfigurable ComputingHW/SW Codesign
HF

Haisheng Fu

The University of British Columbia, Postdoctoral Fellow
Deep LearingImage CompressionVideo CompressionIC Design,Hardware Implementation,Cryptography