AutoGUIWorld: Image Generators as Visual World Models for GUI Agent

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high cost and limited diversity inherent in training GUI agents on real-world environments by pioneering the use of an image generator as a visual world model to synthesize high-quality interaction trajectories without executing actual software. The proposed approach integrates the visual priors of generative models with the task knowledge of planners, achieving data synthesis through structured scene sampling, atomic action planning, iterative screenshot editing, and spatial annotation filtering. In total, 79,000 synthetic samples are generated. Following fine-tuning on this dataset, the resulting model achieves an average score of 40.8% on OSWorld and a success rate of 32.2% on ScienceBoard, demonstrating the efficacy of synthesizing scalable and diverse training data for GUI agent development.
📝 Abstract
GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments. Expanding this coverage requires deploying increasingly diverse and complex software, with specialized applications imposing additional installation, configuration, and runtime costs. We introduce AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. AutoGUIWorld samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, and generates tasks conditioned on those scenes. A planner then specifies atomic actions and their intended visual consequences, while an image generator iteratively edits the current screenshot to produce subsequent observations. Action grounding and transition-level quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and the task success rate on ScienceBoard from 14.0% to 32.2%. These results show that generated trajectories improve GUI-agent performance on real desktop and scientific tasks.
Problem

Research questions and friction points this paper is trying to address.

GUI agent
interaction trajectories
data diversity
visual world model
Innovation

Methods, ideas, or system contributions that make the work stand out.

GUI Agent
Visual World Models
Data Generation Framework
Image Generators
Trajectory Synthesis
💼 Related Jobs
fetch failed
C
Cheng Yang
Hunyuan AI Data Team
Y
Yifan Wu
Hunyuan AI Data Team
Y
Yutao Huang
Hunyuan AI Data Team
Z
Zhaohua Zhang
Hunyuan AI Data Team
Beiduo Chen
Beiduo Chen
ELLIS PhD Student, Ludwig-Maximilians-Universität München
LinguisticsNatural Language Processing
M
Muxi Chen
Hunyuan AI Data Team
C
Chenchen Zhao
Hunyuan AI Data Team
H
Hexuan Deng
Hunyuan AI Data Team
Haolin Yang
Haolin Yang
University of Chicago
large language modelsnatural language processing
G
Geyuan Zhu
Hunyuan AI Data Team
S
Sa Zhu
Hunyuan AI Data Team
Jianhuan Zhuo
Jianhuan Zhuo
Institute of Information Engineering, Chinese Academy of Sciences
Representation LearningRecommendation System
Q
Qiuyong Xiao
Hunyuan AI Data Team
J
Jianhao Ruan
Hunyuan AI Data Team
Y
Yiran Peng
Hunyuan AI Data Team
J
Jiayi Zhang
Hunyuan AI Data Team
T
Tian Ye
Hunyuan AI Data Team
Xinlei Yu
Xinlei Yu
Beijing University of Posts and Telecommunications
Stochastic Geometry
Tianwen Jiang
Tianwen Jiang
Harbin Institute of Technology
Knowledge GraphInformation ExtractionNatural Language Processing
J
Jihong Zhang
Hunyuan AI Data Team
Yuyu Luo
Yuyu Luo
Assistant Professor, HKUST(GZ) / HKUST
Data AgentsLLM AgentsDatabaseText-to-SQLData-centric AI