TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决移动设备上大型语言模型长上下文需求导致的内存瓶颈问题,提出了一种基于预测多级缓存优化的框架TierKV,通过预分配不同级别的缓存来减少内存占用并提高处理速度。
📝 Abstract
Large language models (LLMs) are moving onto mobile devices for increasingly diverse workloads over text, images, video, and audio. These applications often require long contexts, making the Key-Value (KV) cache a dominant memory bottleneck because it grows linearly with sequence length and is accessed at every decoding step. Prior work reduces KV-cache footprint through low-rank compression, token eviction, or flash offloading, but the resulting reconstruction overhead, irreversible token loss, or I/O stalls can offset the benefit of saving memory. We present TierKV, a mobile LLM inference framework built on Predictive Multi-Tier Cache Optimization (PMCO). Before decoding starts, PMCO predicts future cache demand from prefill hidden states and jointly assigns tokens to exact, low-rank, and flash-offloaded tiers under the device memory and accuracy budgets. This formulation retains access to the full context, removes the circular dependency of reactive eviction, and admits a closed-form solver that selects tier boundaries and per-layer ranks at runtime. Across eight text, vision, and audio models on three mobile SoCs, TierKV improves prefill throughput by up to 17.6x over existing mobile LLM frameworks, reduces RAM-resident KV cache by 12.5-34%, thereby enabling substantially longer contexts under the same memory budget, while incurring only minor accuracy degradation.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Mobile Devices
Long Contexts
KV Cache
Memory Bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Predictive Multi-Tier Cache Optimization
PMCO
TierKV
Long-Context On-Device LLMs