Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing visual red teaming attacks, which target only model inference and fail to achieve end-to-end control from visual grounding to browser execution. To this end, this work proposes WebMirage, a framework that reformulates red teaming as an end-to-end "ground-and-execute" problem. Specifically, the method leverages large vision-language models to generate localized visual perturbations, integrating role-slot abstraction, webpage reassembly, and data flow analysis techniques to induce web agents into selecting malicious elements and triggering target actions. Experimental results demonstrate that WebMirage achieves an average attack success rate of 91.9% across various configurations, significantly outperforming baseline methods. Furthermore, the proposed framework effectively circumvents three mainstream agent-level defense mechanisms, highlighting its potency in evaluating and exposing vulnerabilities within autonomous web agents.
📝 Abstract
Modern web agents built on large vision-language models process webpages, select relevant UI elements, and translate model outputs into browser actions. Existing visual red-teaming approaches use adversarial visual content to manipulate this process. However, they primarily target model inference and do not explicitly account for structured input processing or action post-processing. Consequently, model-level success does not establish control over browser execution and cannot reliably characterize end-to-end agent robustness. To address this gap, we formulate red teaming for vision-grounded web agents as an end-to-end grounding-to-execution problem, and introduce WebMirage, a framework that crafts localized visual perturbations that cause agents to select attacker-controlled content and execute the corresponding browser action across varying webpage renderings. It uses a role-slot abstraction and webpage recomposition to capture competition among webpage elements, and dataflow analysis to align optimization with action post-processing. We evaluate WebMirage across four agent configurations and six VLM backbones on 2,250 tasks covering 13 public websites and a sandbox benchmark. WebMirage achieves an average attack success rate of 91.9%, compared with 17.4% for the strongest baseline, and remains effective against three agent-level defenses.
Problem

Research questions and friction points this paper is trying to address.

Web Agents
Adversarial Robustness
Red Teaming
Visual Grounding
Vision-Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Web Agents
Adversarial Visual Perturbations
End-to-End Red Teaming
Visual Grounding
Dataflow Analysis