🤖 AI Summary
This study addresses the failure of context-based safety defenses in multi-step AI agents caused by a lack of state awareness. To overcome this limitation, we propose a stateful policy enforcement engine based on extended regular expressions. By introducing stateful predicates, deferred policy generation, and scoped semantic checking, this method models safety constraints as permissible sequences of tool invocations, thereby enabling dynamic, fine-grained control over agent behavior. Experimental evaluations demonstrate that the proposed engine successfully blocks 93–95% of AgentDojo attacks and 62–85% of Toolathlon attacks while maintaining high task utility. These results indicate a significant improvement in the robustness of multi-step agent systems against adversarial threats.
📝 Abstract
Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing stateful contextual policies. A Sapien policy specifies permitted tool-call sequences using a regular expression extended with stateful predicates, deferred policy generation, and scoped semantic checks. We show that Sapien stays within a few percent of an unconstrained agent's utility. Even if the agent is fully hijacked, Sapien's policies rule out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon (twice as many as tool allowlists on long-horizon tasks).