Receiver-Conditioned Latent Communication gives 94% CacheBack

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck of excessive memory and context overhead in KV cache transmission within multi-agent systems, which frequently exceeds GPU resource limits. To this end, it proposes a "receiver-conditioned communication" paradigm and introduces CacheBack, a training-free method that filters and compresses sender-side KV caches based on attention weight analysis. By transmitting only the information required by the receiver, CacheBack enables precise, on-demand communication across heterogeneous architectures such as Transformers and Mamba. Experimental results demonstrate that the proposed approach eliminates 75% of redundant states, improves accuracy by 14.7%, and reduces latency by 3.2×, while exhibiting consistent generalization across multiple model families.
📝 Abstract
Multi-agent systems distribute large contexts across agents that communicate to solve a task. Text messages are compact but require decoding and may omit evidence the receiving agent needs. Recent latent communication instead transfers KV caches. This avoids text generation and can improve accuracy and latency. However, a full KV cache grows linearly with both the context an individual agent processes, and the number of agents that coordinate together. This raises memory and context costs, often far exceeding available GPU resources and context window sizes. Our key observation is that agents need only send what the receiving agent requires for its local task -- which we call receiver-conditioned communication. The receiver agent passes the sender a small description of its information needs, which serves to filter and compress the sender agent's KV cache. CacheBack is a simple, robust, training-free instance of receiver conditioning based on the sender's attention weights. On FanOutQA, CacheBack with Qwen 3 removes 75% of the state the agent would otherwise receive, improving accuracy by 14.7 percentage points and reducing median task-completion latency by 3.2x relative to text communication. We show comparable improvements across model families that span dense Transformers, Mamba-attention hybrids, and sliding-window attention.
Problem

Research questions and friction points this paper is trying to address.

multi-agent systems
latent communication
KV cache
memory overhead
context window
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Communication
KV Cache Compression
Multi-agent Systems
Training-free
Receiver-Conditioned
🔎 Similar Papers
No similar papers found.