🤖 AI Summary
This study addresses the challenges of sparse rewards and the absence of intermediate action supervision in the online training of computer-use agents. To this end, it proposes Online Policy Self-Distillation (OPSD), a method that leverages real-time graphical user interface (GUI) feedback. Specifically, OPSD exploits execution feedback to generate step-level guidance signals and value scores, jointly optimizing token-level distillation with trajectory-level reinforcement learning via Group Relative Policy Optimization (GRPO). Furthermore, an asynchronous training framework is designed to enhance learning efficiency. Experimental results demonstrate that the proposed approach yields significant performance improvements on the OSWorld benchmark while validating the model's generalization capabilities in out-of-distribution scenarios.
📝 Abstract
Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.