🤖 AI Summary
This study addresses the cross-agent redundancy and context bloat arising from directly forwarding full latent states in multi-agent latent collaboration. We propose LatCom, a framework that introduces a novel cross-agent joint compression mechanism to map multiple agents' latent states into a fixed number of receiver-readable slots, optimizing for task utility rather than latent reconstruction. Furthermore, we design a two-stage training strategy comprising single-sender readability learning followed by multi-sender fusion learning, enabling end-to-end optimization with a frozen receiver. Evaluations on Qwen3-4B demonstrate that, compared to LatentMAS, LatCom achieves an average 2.46× inference speedup and reduces output tokens by 70.3% while maintaining comparable accuracy.
📝 Abstract
LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes the receiver-side context scale with both the number of agents and the reasoning length, increasing computation, memory usage, and collaboration latency. A natural solution is latent compression. But we find that cross-agent redundancy remains unresolved in existing latent compression approaches, which typically compress each sender independently and then concatenate the results. We propose LatCom, a cross-agent latent compression framework for efficient multi-agent latent collaboration. LatCom maps multiple sender latents into a fixed number of receiver-readable and task-relevant slots. Rather than reconstructing all sender hidden states, it optimizes the compressed latents for receiver-side task utility. LatCom trains the compressor in two stages: single-sender readability learning first establishes a latent interface interpretable by the frozen receiver, and multi-sender fusion learning then trains the compressor to fuse complementary evidence and remove redundancy across agents. Experiments on multiple benchmarks with Qwen3-4B show that LatCom achieves an average 2.46x inference speed-up over LatentMAS and reduces output token usage by 70.3% while maintaining comparable average accuracy.