LatCom: Cross-Agent Latent Compression for Efficient Multi-Agent Collaboration

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the cross-agent redundancy and context bloat arising from directly forwarding full latent states in multi-agent latent collaboration. We propose LatCom, a framework that introduces a novel cross-agent joint compression mechanism to map multiple agents' latent states into a fixed number of receiver-readable slots, optimizing for task utility rather than latent reconstruction. Furthermore, we design a two-stage training strategy comprising single-sender readability learning followed by multi-sender fusion learning, enabling end-to-end optimization with a frozen receiver. Evaluations on Qwen3-4B demonstrate that, compared to LatentMAS, LatCom achieves an average 2.46× inference speedup and reduces output tokens by 70.3% while maintaining comparable accuracy.
📝 Abstract
LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes the receiver-side context scale with both the number of agents and the reasoning length, increasing computation, memory usage, and collaboration latency. A natural solution is latent compression. But we find that cross-agent redundancy remains unresolved in existing latent compression approaches, which typically compress each sender independently and then concatenate the results. We propose LatCom, a cross-agent latent compression framework for efficient multi-agent latent collaboration. LatCom maps multiple sender latents into a fixed number of receiver-readable and task-relevant slots. Rather than reconstructing all sender hidden states, it optimizes the compressed latents for receiver-side task utility. LatCom trains the compressor in two stages: single-sender readability learning first establishes a latent interface interpretable by the frozen receiver, and multi-sender fusion learning then trains the compressor to fuse complementary evidence and remove redundancy across agents. Experiments on multiple benchmarks with Qwen3-4B show that LatCom achieves an average 2.46x inference speed-up over LatentMAS and reduces output token usage by 70.3% while maintaining comparable average accuracy.
Problem

Research questions and friction points this paper is trying to address.

multi-agent systems
latent compression
cross-agent redundancy
large language models
efficient collaboration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Agent Latent Compression
Multi-Agent Collaboration
Latent Communication
Two-Stage Training
Task Utility Optimization
🔎 Similar Papers
No similar papers found.
S
Shinan Zhang
University of Science and Technology of China
T
Tao Zhang
University of Science and Technology of China
Q
Qihui Zhu
University of Science and Technology of China
M
Mengjie Zhang
University of Science and Technology of China
D
Dong Jin
University of Science and Technology of China
Y
Yunpeng Hou
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center
S
Shuangwu Chen
University of Science and Technology of China
X
Xiaobin Tan
University of Science and Technology of China
Quan Zheng
Quan Zheng
Institute of Software, Chinese Academy of Sciences
Computer Graphics
J
Jian Yang
University of Science and Technology of China