AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of web agents to adaptive prompt injection attacks and the limited generalizability of existing defenses. We propose a tripartite co-evolutionary framework grounded in a frozen web world model. Through reinforcement learning and curriculum learning, this framework jointly evolves task curricula, adversarial injection generators, and agent policies. A "success-flipping" reward mechanism is introduced to optimize adversarial examples, overcoming the limitations of static training to simultaneously enhance capability and robustness. Experimental results demonstrate that a 4B-parameter agent trained under this framework achieves a 33.6% improvement in task completion rate against unseen frontend-model attacks compared to baselines, while exhibiting zero-shot transferability to real-world browsers.
📝 Abstract
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6\% relative to the base agent.
Problem

Research questions and friction points this paper is trying to address.

Web agents
Prompt injection
Adaptive attacks
Robustness
Adversarial training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Prompt Injection
Web World Model
Co-evolutionary Training
Sim-to-Real Transfer
Curriculum Learning