KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of redundant KV cache computation and excessive memory overhead caused by prefix discrepancies when multiple agents share context. To mitigate this, we propose an online KV cache rectification framework that leverages compact low-rank state representations to capture cross-agent deviations. This approach facilitates dynamic context sharing and seamless chaining along agent workflows, eliminating reliance on reference caches without requiring reference prefilling. Empirically, our method preserves the cache fidelity of the initial agent while matching or surpassing existing performance baselines. Furthermore, it achieves a 2.0× acceleration in time-to-first-token (TTFT) and a 3.7× reduction in peak GPU memory consumption, establishing a new paradigm for efficient and scalable multi-agent context sharing.
📝 Abstract
Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However, agent-specific prefixes change the KV cache generated for the same shared context, causing each agent to repeatedly prefill the growing context and construct a separate cache with high computation and memory overhead. Selective recomputation reduces this redundancy but still retains substantial model execution, while existing delta correction methods either support only recurring context relations or maintain memory-intensive online correction states for dynamically changing context. For first seen shared context, these methods also construct a reference cache outside the agent workflow, and an approximate correction at the first agent affects the outputs passed to subsequent agents. We present KVCMAS, an online KV cache correction framework that represents cross-agent cache deviations using compact low-rank states and seamlessly chains corrections along the agent workflow without an additional reference prefill. This design supports dynamically changing shared context while preserving an exact first-agent cache. Across multiple language and vision-language workloads, KVCMAS matches or improves the accuracy of prior KV cache sharing methods while achieving the lowest TTFT under highly concurrent serving. Under controlled serving traces, it provides a 2.0x TTFT speedup over inference without KV cache sharing and reduces peak GPU memory by up to 3.7x relative to a prior KV cache correction method. These results establish KVCMAS as an accurate and scalable KV cache sharing approach for prompt-specialized multi-agent serving.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Systems
KV Cache Sharing
Prompt-specialized Agents
Computation Overhead
Memory Overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV cache correction
multi-agent systems
low-rank states
shared context
online framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.