Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the vulnerability of LLM-based web agents to unauthorized consequences triggered by legitimate actions within deceptive interfaces. To mitigate this, we propose Veer, a runtime defense framework that establishes task-relevant webpage states as independent control objectives. By proactively modifying environmental states through predictive trajectory intervention, Veer circumvents risks without disrupting the agent’s original task planning. This work is the first to identify the failure mode wherein valid actions induce unauthorized outcomes, achieving a paradigm shift from behavioral interception to state intervention. Furthermore, it ensures intervention reliability by integrating prospective reasoning with temporal evidence analysis. Evaluated on the TrickyArena and WebDecept benchmarks, Veer achieves state-of-the-art safe task completion rates, outperforming the second-best method by 15.9%–25.0% while reducing dark pattern success rates to 0.3%.
πŸ“ Abstract
LLM-based Web agents can autonomously complete user tasks, yet deceptive interfaces can steer them toward outcomes that conflict with users'interests. Existing defenses primarily intervene on agent behavior through blocking, guidance, or replanning. We identify a distinct failure mode: a task-valid action can still realize an unauthorized consequence because of the current Web state. This motivates treating task-relevant Web state itself as a runtime control target. We introduce Veer, an agent-side runtime defense that leaves task planning to the base agent and intervenes on Web state when a proposed action would produce an unauthorized consequence. Before modifying the live environment, Veer constructs a prospective intervention trajectory toward a safe task-relevant state and executes it with runtime grounding and verification. Across TrickyArena and WebDecept, Veer achieves the highest safe task completion in all three evaluation settings, exceeding the next-best defense by 15.9 and 25.0 percentage points on TrickyArena-Single and TrickyArena-Multi, respectively, while reducing dark-pattern success on WebDecept to 0.3%. These gains persist across dark-pattern types and all 12 agent, model, and benchmark configurations. Ablations show that active state intervention provides the largest gain, while prospective rollout and temporal evidence contribute additional improvements. These results establish task-relevant Web state as an effective runtime control target for protecting Web agents from deceptive outcomes.
Problem

Research questions and friction points this paper is trying to address.

Web agents
deceptive interfaces
dark patterns
runtime defense
state intervention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prospective State Intervention
Web Agents
Deceptive Interfaces
Runtime Defense
Dark Patterns
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.