When Collaboration Becomes a Trigger: Collective Evidence-Threshold Backdoors in Multi-Agent Systems

📅 2026-08-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work reveals a novel backdoor threat in multi-agent systems wherein collaborative mechanisms can be exploited to trigger malicious behavior once collectively accumulated evidence among agents surpasses a hidden threshold—a risk this paper terms “collective evidence threshold” backdoors. To demonstrate this vulnerability, the authors introduce the BCBI attack, which leverages counterfactual boundary pairs to achieve highly selective activation with minimal false triggers. In response, they propose LATTE, an unsupervised defense framework that operates solely on clean test-time data by modeling latent state dynamics and detecting anomalous transitions in hidden states. LATTE effectively blocks malicious propagation under unknown attacks while preserving normal system functionality.
📝 Abstract
LLM-based multi-agent systems (MAS) extend LLM capabilities through iterative communication and shared contexts. However, this collaboration introduces a vulnerability: backdoor behavior can be activated when peer evidence reaches a hidden threshold, rather than being determined by any single message. We introduce a collective evidence-threshold backdoor paradigm for MAS and Boundary-Conditioned Backdoor Injection (BCBI), which constructs counterfactual boundary pairs to separate benign behavior before the threshold from the adversarial objective after it, and learns latent progression aligned with evidence. To mitigate this threat, we propose LAtent Transition Test-time Evaluation (LATTE), a clean-only latent-transition defense that learns benign communication dynamics and quarantines anomalous agent updates before their responses propagate. Across several benchmarks, BCBI yields selective activation with little premature activation; without knowing the attack target or trigger, LATTE limits propagation with minimal disruption.
Problem

Research questions and friction points this paper is trying to address.

multi-agent systems
backdoor attack
collective evidence threshold
LLM security
adversarial behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

collective evidence-threshold backdoor
multi-agent systems
backdoor attack
latent transition
test-time defense