StepKV: Step-Aware KV Cache Compression for LLM Agents

📅 2026-08-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决多步LLM代理中KV缓存压缩问题,StepKV通过考虑推理步骤和令牌级别信息,保留关键推理步骤,提高效率与准确性。
📝 Abstract
Key-value (KV) caching is essential for efficient autoregressive large language model (LLM) inference, but the cache grows linearly with context length, increasing storage and decoding costs. KV cache compression mitigates this cost by retaining only a subset of cached tokens. This challenge is particularly important for multi-step LLM agents, where a query expands into trajectories of reasoning, tool interactions, and retrieved observations. Existing pruning methods typically treat the cache as a flat token stream and rank tokens by recency or attention saliency. This creates a mismatch between the unit of compression and the unit of reasoning: token-level pruning removes individual entries, whereas useful information in multi-step agents is often organized into reasoning steps with uneven and delayed importance. Consequently, an early observation or intermediate decision may receive little recent attention yet remain essential for later evidence synthesis. We term this failure mode Reasoning Continuity Disruption.These observations motivate KV cache compression that jointly considers token- and reasoning-step-level information. StepKV addresses this goal by treating reasoning steps as first-class retention units. It associates cache entries with their generating steps, estimates step utility from trajectory-derived signals, and combines this utility with token-level saliency. The resulting scores globally rank prunable tokens, from which StepKV retains the top-scoring entries under a target budget. StepKV thus provides a step-centric perspective for agent KV cache compression. Across multi-hop QA and long-horizon web reasoning tasks, StepKV sustains accuracy under low KV budgets where token-level baselines degrade sharply, offering a more robust efficiency-accuracy trade-off for multi-step agent inference.
Problem

Research questions and friction points this paper is trying to address.

KV Cache Compression
Multi-step LLM Agents
Reasoning Continuity Disruption
Innovation

Methods, ideas, or system contributions that make the work stand out.

Step-Aware
KV Cache Compression
Reasoning Continuity Disruption
Multi-Step LLM Agents
B
Boyu Feng
The Chinese University of Hong Kong
J
Jiahong Liu
The Chinese University of Hong Kong
Y
Yifan Li
The Chinese University of Hong Kong
W
Wenhao Yu
The Chinese University of Hong Kong
Zexuan Qiu
Zexuan Qiu
The Chinese University of Hong Kong
Natural Language Processing
Y
Yuliang Sun
The Chinese University of Hong Kong
M
Ming Shen
The Chinese University of Hong Kong
X
Xiang Li
Huawei Technologies Co., Ltd
Q
Quanyu Dai
Huawei Technologies Co., Ltd
Irwin King
Irwin King
The Chinese University of Hong Kong
social computingmachine learningAIgraph neural networksNLP