Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high costs and auditability challenges arising from the isolated evaluation of defense mechanisms in multi-agent systems by proposing the DEFER framework. DEFER constructs a cascaded defense architecture that innovatively delineates the boundary between deterministic rules and evaluative judgments. It employs a hybrid approach wherein deterministic checks are prioritized, while residual cases are delegated to a panel of multiple judges. Experimental results demonstrate that this method reduces the attack success rate from 30% to 3%, with 78% of attacks intercepted by deterministic rules. Consequently, DEFER substantially enhances system security while significantly reducing computational overhead.
📝 Abstract
LLM-based multi-agent systems (MAS) engage tools, share memory, and delegate tasks, often encountering adversarial content. Current defenses for MAS are typically evaluated in isolation, focusing on one attack type at a time, which can lead to costly and hard-to-audit outcomes. This study organizes defenses into five principles, implementing them as DEFER1 (DEterministic-First Enforcement with Residual judgment), which includes a cascade of 28 checks that blocks what it can and refers the rest to a panel of four judges. In independent testing across four domains, attack success rates drop from about 30.0% to approximately 3.0%, with 78% of blocked attacks handled by deterministic checks. Only a quarter of proposals reach the judges in the security-operations domain, illustrating that the rules provide security for attacks violating clear policies, while judges manage those that only misrepresent intent. Both systems have weaknesses, such as a risk-score approval gate that inaccurately approves most attack proposals but few legitimate ones, highlighting the challenges in assessing threats accurately.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Systems Security
Adversarial Content
Defense Evaluation
Judgment Boundary
Threat Assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Systems Security
DEFER1
Deterministic-First Enforcement
LLM Judges
Cascade Defense
🔎 Similar Papers