OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing AI safety evaluations in capturing security risks arising from cumulative changes in long-horizon, statefully evolving environments. To this end, the authors propose the first red-teaming framework capable of supporting large-scale state evolution and tool interaction, introducing Evolutionary Markov Hypergraph Attacks (EMHA)—a model-parameter-free environmental perturbation strategy. Leveraging a skill repository of over 500,000 tools, EMHA generates more than 10,000 state-dependent, long-horizon task scenarios across 50 domains to uniformly evaluate 75 agent configurations. Experimental results demonstrate that EMHA achieves an average attack success rate of 85.0%, outperforming baseline methods relying solely on instruction evolution by over 17% in complex environments, thereby effectively uncovering the critical impact of runtime environmental dynamics on agent safety.
📝 Abstract
AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety benchmarks often fail to capture these cumulative risks because they focus on short, static tasks. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART provides over 10,000 validated stateful scenarios across 50 domains, drawing from a pool of more than 500,000 tools and skills. These tasks require a median of 97 tool calls and enable unified evaluation across 75 different agent-model configurations. To systematically explore these evolving attack surfaces, we propose the Evolutionary Markov Hypergraph Attack (EMHA). EMHA is a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates. Throughout the evaluation, task objectives remain fixed while only the environment state changes. Across all configurations, EMHA achieves a pooled Attack Success Rate (ASR) of 85.0%. Its advantage over instruction-only evolution increases from approximately 2% on simple environments to over 17% on the most complex ones, demonstrating that environment evolution increasingly exposes safety failures as task complexity grows. Furthermore, our analysis shows that the specific runtime implementation of an agent explains a significant portion of safety variation beyond the underlying model's capabilities. These results establish OpenART as a scalable foundation for studying agent safety in complex, evolving environments.
Problem

Research questions and friction points this paper is trying to address.

agent safety
red teaming
environment evolution
stateful scenarios
long-horizon workflows
Innovation

Methods, ideas, or system contributions that make the work stand out.

agent red teaming
environment evolution
stateful scenarios
Evolutionary Markov Hypergraph Attack
scalable safety evaluation
🔎 Similar Papers
No similar papers found.