🤖 AI Summary
This work addresses the limitations of existing AI safety evaluations in capturing security risks arising from cumulative changes in long-horizon, statefully evolving environments. To this end, the authors propose the first red-teaming framework capable of supporting large-scale state evolution and tool interaction, introducing Evolutionary Markov Hypergraph Attacks (EMHA)—a model-parameter-free environmental perturbation strategy. Leveraging a skill repository of over 500,000 tools, EMHA generates more than 10,000 state-dependent, long-horizon task scenarios across 50 domains to uniformly evaluate 75 agent configurations. Experimental results demonstrate that EMHA achieves an average attack success rate of 85.0%, outperforming baseline methods relying solely on instruction evolution by over 17% in complex environments, thereby effectively uncovering the critical impact of runtime environmental dynamics on agent safety.
📝 Abstract
AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety benchmarks often fail to capture these cumulative risks because they focus on short, static tasks. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART provides over 10,000 validated stateful scenarios across 50 domains, drawing from a pool of more than 500,000 tools and skills. These tasks require a median of 97 tool calls and enable unified evaluation across 75 different agent-model configurations. To systematically explore these evolving attack surfaces, we propose the Evolutionary Markov Hypergraph Attack (EMHA). EMHA is a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates. Throughout the evaluation, task objectives remain fixed while only the environment state changes. Across all configurations, EMHA achieves a pooled Attack Success Rate (ASR) of 85.0%. Its advantage over instruction-only evolution increases from approximately 2% on simple environments to over 17% on the most complex ones, demonstrating that environment evolution increasingly exposes safety failures as task complexity grows. Furthermore, our analysis shows that the specific runtime implementation of an agent explains a significant portion of safety variation beyond the underlying model's capabilities. These results establish OpenART as a scalable foundation for studying agent safety in complex, evolving environments.