SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the covert attack-defense challenges in tool-use agents arising from state changes by proposing a unified framework based on Partially Observable Markov Decision Processes (POMDPs). On the offensive side, the DART algorithm is introduced to decompose execution steps and dynamically search trajectories. On the defensive side, the SAGE mechanism is developed to actively verify environmental states through read-only queries, enabling precise interception of malicious actions. Additionally, an environment-verifiable dataset is constructed for evaluation. Experimental results demonstrate that DART improves the attack success rate by 18.8%–35.9%, while SAGE successfully intercepts 92.7% of harmful trajectories, reducing the online attack success rate to merely 4%.
📝 Abstract
Language-model agents increasingly use tools to act on external systems. Earlier actions can alter files, permissions, database records, or other state, making a later routine-looking action harmful. Yet the visible interaction may not reveal the underlying state needed to assess that action. We formulate attack and defense as partially observed state control in SEAD, deriving their design requirements from this shared execution process. Because attackers supply instructions while the target chooses concrete actions, DART decomposes harmful goals into locally plausible steps and uses feedback from actual tool execution to guide trajectory search. The defender must decide before execution with incomplete state evidence. SAGE can therefore investigate relevant state through read-only queries before allowing or blocking each action, including those proposed after a block. We construct an environment-verifiable dataset integrating controlled initial states, replayable tool environments, and task-specific executable checks. Across four target models, DART improves semantic attack success by 18.8--35.9 percentage points over the competing baseline, with consistent gains under executable verification. On recorded trajectories, SAGE preserves 95.79% of benign trajectories while intercepting 92.73% of harmful paths by the harm-enabling boundary. In online attack-defense evaluation, it reduces DART's executable attack success from 48.0% to 4.0%. SAGE remains effective across four attack methods and generalizes to out-of-domain environments. Our code and data is available at https://github.com/EverywhereSafety/SEAD.
Problem

Research questions and friction points this paper is trying to address.

Tool-Using Agents
State-Based Security
Attack and Defense
Partially Observed State
LLM Agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tool-Using Agents
Partially Observed State Control
Adversarial Attack and Defense
Trajectory Search
State Probing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.