Imagine to Act: High-Fidelity Data Synthesis via Image Editing World Model for Scalable GUI Agent Training

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of high-fidelity, scalable visual action trajectory data for training GUI agents by proposing the Infinite-Dreamer framework. This approach introduces a pioneering pixel-level image editing-based world model that formulates GUI state transitions as an image editing task. Specifically, it leverages vision-language models to generate structured delta text descriptions of interface changes and fine-tunes an image editing backbone to synthesize high-fidelity screenshot transitions and multi-step imagined trajectories. This enables simulator-free data synthesis while precisely preserving critical visual details such as icon layouts. Experimental results demonstrate that the proposed method significantly outperforms existing baselines on benchmarks including AndroidWorld, improving Pass@1 by 4.45% and 9.05% for 8B and 2B models, respectively.
📝 Abstract
Graphical User Interface (GUI) agents have emerged as a promising paradigm for automating complex digital workflows across diverse applications. However, training highly capable and generalizable agents fundamentally relies on massive, high-fidelity visual-action trajectories, which are notoriously difficult to acquire. While human demonstrations are unscalable, existing GUI world models rely on text descriptions or HTML rendering, discarding crucial pixel-level visual details like icons and layout styles. To address this issue, we introduce Infinite-Dreamer, a simulation-free data synthesis method powered by a pixel-level Image Editing World Model. By conceptualizing GUI transitions as image editing tasks, we leverage Vision-Language Models (VLMs) to describe action-induced UI changes as structured delta-text. We then fine-tune an image editing backbone to controllably synthesize realistic screenshot transitions. We utilize this model to generate both single-frame visual robustness data and multi-step imaginary trajectories. To validate the effectiveness of our approach, we fine-tune the Qwen3-VL baseline solely on the synthesized data to obtain Infinite-Actor, and evaluate it on AndroidWorld, MobileWorld, and AndroidControl-Curated benchmarks. Infinite-Actor consistently outperforms the Qwen3-VL baselines across scales: Infinite-Actor-8B improves AndroidWorld Pass@1 by +4.45 and nearly doubles the MobileWorld Pass@3 success rate, while Infinite-Actor-2B improves Pass@1 by +9.05. Code is available at https://github.com/swaydy-n/Infinite-Dreamer.
Problem

Research questions and friction points this paper is trying to address.

GUI Agent
Data Synthesis
World Model
Visual-Action Trajectories
High-Fidelity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Image Editing World Model
Simulation-free Data Synthesis
GUI Agent
Vision-Language Models
Imaginary Trajectories