SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems

📅 2026-07-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the security challenge in multi-agent systems where malicious objectives can evade single-point detection by decomposing tasks, thereby triggering information leakage or hazardous operations. To counter this threat, the paper introduces semantic information flow control into multi-agent security for the first time, proposing a framework that integrates structured semantic taint labeling, dynamic collaboration graph propagation, and workflow-level semantic validation. This approach reconstructs global risk context prior to irreversible actions, enabling cross-delegation-boundary tracking of malicious intent while preserving risk visibility. Empirical evaluation across four benchmark scenarios demonstrates that the method significantly reduces attack success rates without compromising the completion rate of benign tasks, while also exhibiting strong discriminative capability between safe and harmful behaviors.
📝 Abstract
Multi-agent systems improve capability through task decomposition and role specialization, but these same mechanisms introduce an important safety blind spot: a harmful objective can be fragmented into locally plausible subtasks, allowing malicious intent to evade detection by any single agent. This is a growing social-impact challenge: systems handling sensitive information or consequential tools can turn routine delegation into unauthorized disclosure or unsafe action. We argue that this failure mode is better understood as a semantic information-flow problem than as a single-turn prompt classification task. To address this, we propose SafeFlow, a defense framework for multi-agent systems that formalizes malicious cross-agent propagation as a semantic information-flow problem. SafeFlow attaches structured semantic taints to root requests, propagates them through a dynamic collaboration graph, and performs workflow-level validation to reconstruct the global risk context before irreversible actions are committed. Evaluated on four benchmarks spanning prompt injection, jailbreak-based unsafe tool use, risky code execution, and harmful web-agent behavior, SafeFlow reduces attack success rates compared to undefended baselines and external defenses while retaining high benign task completion and a high paired safe--harm success rate. Our findings show that multi-agent systems still lack mechanisms for preserving risk semantics across delegation boundaries. This gap can turn routine delegation into privacy harms or unsafe actions that affect people and organizations. SafeFlow keeps this risk visible throughout the workflow, before it results in harm.
Problem

Research questions and friction points this paper is trying to address.

multi-agent systems
semantic information-flow
malicious propagation
task delegation
safety blind spot
Innovation

Methods, ideas, or system contributions that make the work stand out.

semantic information-flow
multi-agent systems
structured semantic taints
workflow-level validation
malicious propagation
🔎 Similar Papers
No similar papers found.