Safety of Latent Communication in Multi-Agent Systems

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of multi-agent systems in which latent communication channels can circumvent safety alignment to elicit harmful compliance. We reveal that lightweight communication mappings introduce exploitable vulnerabilities, and that even benign training can amplify such security risks. To exploit these weaknesses, we propose a novel reinforcement learning-based attack method that operates without requiring target responses. Furthermore, we introduce a reward-only adjustment mechanism as a corresponding defense. Experimental results demonstrate that our attack increases the harmful compliance rate from 27.9% to 76.9% while preserving high task accuracy. Conversely, the proposed mitigation strategy significantly reduces system harm across diverse attack scenarios. Collectively, this work establishes a new paradigm for achieving holistic safety alignment in multi-agent systems.
📝 Abstract
Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Systems
Latent Communication
Safety Alignment
Adversarial Attack
Harmful Compliance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Communication
Multi-Agent Systems
Reinforcement Learning Attack
Safety Alignment
Link Poisoning