Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference

📅 2026-04-19
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses critical limitations in KV cache management for large-scale GPU inference—namely, fixed sizing, single-level storage, and passive eviction policies—that lead to excessive memory waste and redundant computation. The authors propose a unified KV cache management system that enables precise sizing for diverse attention architectures, including MLA, and introduces a six-tier heterogeneous memory hierarchy spanning from HBM to parallel file systems. A key innovation is a block-type-aware Bayesian reuse prediction model based on Beta conjugate priors, which facilitates head-granularity eviction and RoPE-aware prefetching. Experimental results demonstrate that the system achieves 70–84% cache hit rates, reduces first-token latency by 1.4–2.1×, improves throughput by 1.7–2.9×, lowers inference costs by 47%, and delivers an effective cache capacity of 38 TB per node.
📝 Abstract
Key-value (KV) cache memory management is the primary bottleneck limiting throughput and cost-efficiency in large-scale GPU inference serving. Current systems suffer from three compounding inefficiencies: (1) the absence of unified KV cache sizing across all attention architectures--particularly multi-head latent attention (MLA), which is unsupported in general-purpose frameworks, resulting in up to 57x memory over-provisioning; (2) confinement of KV cache to a single memory tier (GPU HBM) despite the availability of a rich hierarchy spanning CPU DRAM, CXL-attached memory, NVMe via GPUDirect Storage, RDMA fabric, and parallel filesystems; and (3) reactive eviction policies that discard reusable state, forcing redundant recomputation. We present a unified system that addresses all three problems. Our architecture-variant-aware sizing engine computes exact memory requirements per attention type, enabling up to 7.4x higher batch sizes. A six-tier memory hierarchy extends effective KV cache capacity from 40 GB to over 38 TB per node while maintaining sub-millisecond time-to-first-token (TTFT) for hot entries. A Bayesian reuse predictor with Beta conjugate priors over 16 (block-type, transition-type) pairs achieves 70-84% cache hit rates, combined with EMA-scored head-granular eviction and RoPE-aware prefetching. Component-level validation on trace replay using ShareGPT, LMSYS-Chat-1M, and agentic workloads demonstrates 70-84% cache hit rates. Analytical projections combining validated component behavior with published hardware specifications indicate 1.4-2.1x projected TTFT reduction, 1.7-2.9x throughput improvement, and 47% cost reduction compared to state-of-the-art baselines.
Problem

Research questions and friction points this paper is trying to address.

KV cache
memory management
large-scale GPU inference
multi-tier memory
attention architecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV cache management
multi-tier memory hierarchy
Bayesian reuse prediction
attention-variant-aware sizing
RoPE-aware prefetching
🔎 Similar Papers
No similar papers found.