ActKV: Efficient LLM Agents through Action-Guided KV Cache Management

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory bottlenecks and throughput limitations caused by KV cache accumulation during LLM-based agent inference, proposing the first KV cache compression framework tailored for agentic reasoning. The method introduces a novel action-contribution-based eviction criterion that leverages stable access patterns to preserve critical entries. By integrating confidence-driven adaptive budget allocation, page-aware compression primitives, and customized kernel techniques, it achieves efficient coordination between dynamic budgets and paged memory. Experimental results demonstrate that on long-trajectory tasks, the proposed framework retains 98.53% of the original accuracy while utilizing only 26% of peak memory, yielding 3.97× and 3.58× improvements in token and task throughput, respectively.
📝 Abstract
Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output quality, overlooking the asymmetric importance of actions in driving task progress. Our key idea is to establish a compression criterion that values KV entries by their contribution to action generation and prioritizes action quality. However, iterative execution, dynamic memory demands, and scattered action-critical entries pose challenges to eviction policies, budget allocation, and paged memory integration. To this end, we propose ActKV, the first KV cache compression framework tailored for agentic LLM inference. (i) Action-oriented KV cache eviction exploits stable action access patterns to retain entries critical to future actions, supporting reliable task progress under compression. (ii) Confidence-driven adaptive budget allocation uses LLM's intrinsic confidence to adapt the budget to evolving action-critical memory demands. (iii) Page-aware compression management standardizes compression into three primitives with customized kernels, realizing practical throughput gains. On long-trace tasks, ActKV retains an average of 98.53% of FullKV's accuracy with only 25.98% of its peak KV cache memory. It also achieves 3.97 times and 3.58 times FullKV's token and task throughput, delivering state-of-the-art performance.
Problem

Research questions and friction points this paper is trying to address.

KV cache compression
LLM agents
memory overhead
serving throughput
agentic inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV Cache Compression
Agentic LLM Inference
Action-Oriented Eviction
Adaptive Budget Allocation
Page-Aware Management
Z
Zihan Wang
University of Science and Technology of China
C
Cheng Tang
University of Science and Technology of China
L
Lei Gong
University of Science and Technology of China
C
Chao Wang
University of Science and Technology of China
Wenqi Lou
Wenqi Lou
University of Science and Technology of China
FPGA AcceleratorAlgorithm-hardware Co-Optimization
Teng Wang
Teng Wang
University of Science and Technology of China
AcceleratorFPGAArchitecture
X
Xuehai Zhou
University of Science and Technology of China