DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of target localization in complex desktop environments and the scarcity of densely annotated data for computer-use agents. We propose a controllable desktop environment construction framework that dynamically transforms the states and layouts of real-world applications. By integrating screenshots, accessibility trees, and window geometry information, our method automatically generates high-density element annotations, yielding DeskForge-1M, a large-scale dataset comprising 1.2 million samples. Fine-tuning vision-language models on this dataset improves Qwen3.5-4B’s accuracy by over 10% on GUI benchmarks and substantially increases completion rates for long-horizon tasks. This work effectively bridges the critical gap in large-scale, dense supervision data for dynamic desktop scenarios, enabling more robust agent performance.
📝 Abstract
Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, window layout, appearance, and resolution, and fuses screenshots, accessibility trees, and window geometry into dense element annotations while recording the outcome of each executed action. Using this environment, we construct DeskForge-1M, a corpus of 1.2M annotated desktop observations containing 159.7M element instances. We fine-tune four vision-language models on 200K grounding examples drawn from DeskForge-1M. All four improve across held-out desktop conditions and on all five external GUI grounding benchmarks; for Qwen3.5-4B, accuracy increases by 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G. The gains also translate to long-horizon task completion: under a fixed planner, the fine-tuned action models solve more WebArena-Infinity and OpenApps tasks, with Qwen3.5-4B increasing from 31 to 50 of 119 tasks and from 3 to 15 of 100 tasks, respectively. These results show that controllable composition of real desktop environments provides a scalable source of supervision for improving both GUI grounding and long-horizon computer use. The framework code, the dataset, and the fine-tuned model are available from the project page: https://saidgurbuz.github.io/deskforge/
Problem

Research questions and friction points this paper is trying to address.

computer-use agents
GUI grounding
dense supervision
desktop environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Computer-use agents
GUI grounding
Controllable desktop environment
Dense supervision
Vision-language models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.