🤖 AI Summary
This study addresses the challenge of deploying GUI agents, where irreversible actions and the absence of ground-truth labels hinder effective model weight updates from single interactions. To overcome this, we propose SOLO, a framework enabling test-time adaptation under unsupervised signals. Specifically, SOLO introduces an auxiliary judge to evaluate successful trajectories and reconstruct failure prefixes, combined with a proposer-verifier mechanism and Top-K filtering. It further performs real-time self-distillation on lightweight adapters through sliding-window memory management. This work pioneers a minimal weight-space adaptation mechanism relying solely on single interactions. Experiments on benchmarks such as WebArena demonstrate that SOLO improves the success rates of UI-TARS and Qwen3-VL by 3 to 6 percentage points, significantly outperforming existing memory-based approaches.
📝 Abstract
GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI actions can be irreversible, a deployed agent gets one attempt per task occurrence, in arrival order, and every attempt counts. No ground truth is available at any point. We define fully test-time adaptation for GUI agents by these constraints and pair it with a minimal weight-space method, SOLO. Auxiliary models read each episode: a judge selects the episodes it deems successful, and a proposer-verifier pair relabels a failed episode's prefix with the subtask that prefix completed. Admitted episodes enter a short sliding window, and each admission updates a small adapter by top-K self-distillation on the agent's own predictions, provided the window holds a judged success. On recurring task streams built from WebArena, VisualWebArena and MobileWorld, SOLO improves on the frozen agent with both UI-TARS-7B and Qwen3-VL-8B, by three to six points of success rate, and exceeds two in-setting memory methods on the web streams.