Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入GUI-SD-v2,一种两阶段训练框架,改进了多轮GUI交互中的特权指导问题,从而增强了基于图形用户界面代理的任务执行能力。
📝 Abstract
Graphical User Interface (GUI) agents enable the fulfillment of complex user instructions through multi-turn interactions with software environments, requiring step-wise reasoning and long-horizon memory to guide actions and retain task-relevant information, respectively. Recent on-policy self-distillation (OPSD) methods have achieved strong performance on GUI grounding, a foundational subtask for GUI agents, owing to dense token-level supervision from privilege-conditioned self-teachers. However, extending existing OPSD methods to multi-turn GUI agents is hindered by self-teachers' limited privilege-following ability and insufficient privileged guidance. In this paper, we introduce GUI-SD-v2, the next version of GUI-SD, which extends OPSD from GUI grounding to multi-turn GUI interaction and addresses key limitations through a two-stage training framework. Specifically, GUI-SD-v2 first strengthens privilege following by jointly optimizing rollouts with and without privileged guidance from the same GUI states. Furthermore, it selectively distills step-specific reasoning and memory guidance through a privilege-conditioned self-teacher, supporting action decisions and the retention of task-relevant information for subsequent interactions. Extensive experiments on two representative GUI agent benchmarks, AndroidWorld and MobileWorld, show that GUI-SD-v2 compares favorably with existing OPSD baselines while consistently outperforming the evaluated state-of-the-art methods in both Pass@1 and Pass@3 success rates. Code and training data will be publicly released.
Problem

Research questions and friction points this paper is trying to address.

Graphical User Interface (GUI) Agents
Multi-turn Interaction
On-policy Self-distillation (OPSD)
Privileged Guidance
Step-wise Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

on-policy self-distillation
multi-turn GUI interaction
two-stage training framework
privilege-conditioned self-teacher
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yan Zhang
Institute of Information Engineering, Chinese Academy of Sciences
Daiqing Wu
Daiqing Wu
Institute of Information Engineering, CAS
Machine learning
Huawen Shen
Huawen Shen
Phd of Chinese Academy of Science
L
Liang Li
Institute of Information Engineering, Chinese Academy of Sciences
G
Gang Cao
Tencent
Z
Zhi Gong
Tencent
W
Wei Dai
Tencent
X
Xiaode Zhang
Tencent
Can Ma
Can Ma
Unknown affiliation
Y
Yu Zhou
VCIP & TMCC & DISSec, College of Computer Science, Nankai University