🤖 AI Summary
This study addresses the high costs and auditability challenges arising from the isolated evaluation of defense mechanisms in multi-agent systems by proposing the DEFER framework. DEFER constructs a cascaded defense architecture that innovatively delineates the boundary between deterministic rules and evaluative judgments. It employs a hybrid approach wherein deterministic checks are prioritized, while residual cases are delegated to a panel of multiple judges. Experimental results demonstrate that this method reduces the attack success rate from 30% to 3%, with 78% of attacks intercepted by deterministic rules. Consequently, DEFER substantially enhances system security while significantly reducing computational overhead.
📝 Abstract
LLM-based multi-agent systems (MAS) engage tools, share memory, and delegate tasks, often encountering adversarial content. Current defenses for MAS are typically evaluated in isolation, focusing on one attack type at a time, which can lead to costly and hard-to-audit outcomes. This study organizes defenses into five principles, implementing them as DEFER1 (DEterministic-First Enforcement with Residual judgment), which includes a cascade of 28 checks that blocks what it can and refers the rest to a panel of four judges. In independent testing across four domains, attack success rates drop from about 30.0% to approximately 3.0%, with 78% of blocked attacks handled by deterministic checks. Only a quarter of proposals reach the judges in the security-operations domain, illustrating that the rules provide security for attacks violating clear policies, while judges manage those that only misrepresent intent. Both systems have weaknesses, such as a risk-score approval gate that inaccurately approves most attack proposals but few legitimate ones, highlighting the challenges in assessing threats accurately.