ORBIT: A Framework for Multi-Agent Safety and Security Evaluations

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of a unified security evaluation benchmark for multi-agent large language model (LLM) systems by proposing a configurable security assessment framework that supports flexible definitions of communication topologies, agent roles, and attack-defense scenarios. Built upon the Inspect framework and integrated with five categories of real-world tasks, this work presents the first empirical investigation jointly examining attacks, defenses, and system architectures. The results demonstrate that no universally effective defense exists: specific strategies fail under multi-agent collusion, and defenses do not transfer across distinct threat models. Furthermore, this research quantifies the trade-off between security and task performance, providing critical empirical evidence to guide the secure design of multi-agent LLM systems.
📝 Abstract
Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet they give rise to significant safety and security risks: the flexible protocols that enable task generalization also expose novel threats, from cascading prompt injection to inter-agent collusion. Progress in defending against these threats has been slowed by a lack of shared empirical infrastructure, which forces bespoke environment development for every new defense and makes standardized comparison impossible. Existing evaluations address isolated threat models or single-agent settings, but none jointly vary attack, defense, and architecture across realistic multi-agent environments. To address this gap, we introduce ORBIT, a configurable evaluation framework for empirical multi-agent safety and security research, built on UK AISI's Inspect. ORBIT lets researchers configure communication topologies, memory, scheduling, and agent roles. It supports four threat types and four defense strategies, as well as non-adversarial failures, with a benchmark suite spanning five scenario families covering browser use, computer use, agentic coding, customer service, and cooperative allocation. Our central finding is a gap in defense transferability across threats: per-action defenses that cut a compromised agent's attack success by 60 points on multi-issue coding give no measurable protection against colluding agents, and none of the defenses we tested generalized over all attacks tested. We further demonstrate security-performance tradeoffs and interactions between architecture and defense effectiveness. We make ORBIT available open-source at https://github.com/wlanderson0/orbit.
Problem

Research questions and friction points this paper is trying to address.

multi-agent systems
safety and security
prompt injection
evaluation framework
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Systems
Safety Evaluation
Prompt Injection
Defense Transferability
Benchmark Framework