One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of deploying GUI agents, where irreversible actions and the absence of ground-truth labels hinder effective model weight updates from single interactions. To overcome this, we propose SOLO, a framework enabling test-time adaptation under unsupervised signals. Specifically, SOLO introduces an auxiliary judge to evaluate successful trajectories and reconstruct failure prefixes, combined with a proposer-verifier mechanism and Top-K filtering. It further performs real-time self-distillation on lightweight adapters through sliding-window memory management. This work pioneers a minimal weight-space adaptation mechanism relying solely on single interactions. Experiments on benchmarks such as WebArena demonstrate that SOLO improves the success rates of UI-TARS and Qwen3-VL by 3 to 6 percentage points, significantly outperforming existing memory-based approaches.
📝 Abstract
GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI actions can be irreversible, a deployed agent gets one attempt per task occurrence, in arrival order, and every attempt counts. No ground truth is available at any point. We define fully test-time adaptation for GUI agents by these constraints and pair it with a minimal weight-space method, SOLO. Auxiliary models read each episode: a judge selects the episodes it deems successful, and a proposer-verifier pair relabels a failed episode's prefix with the subtask that prefix completed. Admitted episodes enter a short sliding window, and each admission updates a small adapter by top-K self-distillation on the agent's own predictions, provided the window holds a judged success. On recurring task streams built from WebArena, VisualWebArena and MobileWorld, SOLO improves on the frozen agent with both UI-TARS-7B and Qwen3-VL-8B, by three to six points of success rate, and exceeds two in-setting memory methods on the web streams.
Problem

Research questions and friction points this paper is trying to address.

GUI agents
test-time adaptation
online learning
single rollout
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Adaptation
GUI Agents
Self-Distillation
Adapter Tuning
Single Rollout
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ziqiang Wang
Ziqiang Wang
Concordia University
Computer Vision
L
Li Gu
Concordia University, Mila – Québec AI Institute
Zhixiang Chi
Zhixiang Chi
University of Toronto
Computer VisionMachine Learning
L
Linlian Jiang
Concordia University, Mila – Québec AI Institute
Z
Zihuan Jiang
University of Toronto
Linqiang Guo
Linqiang Guo
Concordia University
LLMLVMCVGUI Testing
S
Siobhan Reid
Concordia University, Mila – Québec AI Institute
Z
Zhi Liu
Shanghai University
Yang Wang
Yang Wang
Computer Science, Concordia University
computer visionmachine learningdeep learningartificial intelligence