Capture the lifecycle: KV Cache management in ReAct Agents with KVTether

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficient KV cache management and high latency in ReAct agents caused by semantic loss. To this end, we propose KVTether, a framework that captures runtime states by tracing semantic primitives and maps message-level semantics to KV-level lifetimes. KVTether introduces the first lifetime-aware cache management mechanism, bridging the semantic gap between agents and the underlying serving stack to prevent premature eviction of reusable data. Experimental results demonstrate that, compared with LMCache and MORI, KVTether reduces end-to-end latency by up to 26.3% and 17.4%, respectively, while decreasing average task costs by 40.0% and 33.2%.
📝 Abstract
Efficient serving of long-context reasoning-and-acting (ReAct) agents relies on KV cache reuse to reduce large language model (LLM) prefill latency and monetary cost. However, a semantic gap exists between agent harnesses and the underlying serving stack. Through context mutation, tool execution, and subagent coordination, context messages may become actively engaged, permanently discarded, and temporarily unused, while the serving stack only observes accesses to the corresponding KV cache. This lifecycle blindness prevents recency-only policies such as LRU from reclaiming dead KV promptly and from preserving older KV that will be reused sooner than newer entries. We present KVTether, a lifecycle-aware KV cache management framework for ReAct agents. By tracing semantic primitives embedded in agent harnesses, KVTether captures runtime lifecycle semantics during highly dynamic execution. KVTether then translates message-level semantics into KV-level lifecycle states and uses these states to drive state-prioritized cache management without exposing physical complexities to agent harnesses. After reclaiming dead KV, KVTether preferentially preserves live-but-idle KV that is waiting for reuse, reducing premature eviction before reuse. Across agent benchmarks and production workloads, KVTether reduces end-to-end request latency by up to 26.3% and 17.4% relative to LMCache and MORI, respectively, and lowers estimated task cost by 40.0% and 33.2% on average.
Problem

Research questions and friction points this paper is trying to address.

KV cache management
ReAct agents
lifecycle awareness
semantic gap
cache eviction
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV Cache Management
ReAct Agents
Lifecycle-aware
Semantic Gap
Cache Eviction
🔎 Similar Papers
No similar papers found.