🤖 AI Summary
This study addresses the challenge that GUI agents struggle with complex tasks due to knowledge gaps and prior drift. To this end, we propose a robust training paradigm adapted to real-world environmental shifts. Methodologically, vision-language models are leveraged to construct UI state transition graphs for structured exploration. Unlabeled trajectory synthesis is employed to inject five categories of realistic errors, combined with a noise-aware reinforcement learning strategy that enables agents to acquire reliable behaviors under imperfect priors. Experimental results demonstrate that the proposed approach significantly improves screen element discovery rates and cross-dataset accuracy while effectively rejecting erroneous priors, thereby providing a highly robust solution for complex GUI interactions.
📝 Abstract
GUI agents built on large language and vision-language models still struggle on unseen applications and complex multi-step tasks, as completing real GUI tasks depends on app-specific, temporally volatile operational knowledge that is scarce in pretraining corpora. Retrieval-augmented execution offers a natural remedy but faces two coupled bottlenecks: knowledge at scale is hard to acquire, and self-collected priors inevitably drift from the live environment due to version updates, promotions, ads, A/B tests, and personalization. We therefore argue that GUI agents should not pursue perfect knowledge but learn to act correctly under imperfect priors, and propose our framework that couples knowledge acquisition with noise-robust utilization: a structured exploration strategy traverses interactive elements, builds a UI state-transition graph, and synthesizes (task, trajectory) pairs via a VLM without human annotation; a noise-aware training strategy, grounded in a taxonomy of real GUI drift patterns, injects five types of realistic errors into self-explored trajectories to teach the agent to assess prior reliability before acting. Experiments on physical devices and online emulator benchmarks show that our method discovers more unique screens, covers more benchmark tasks, and more effectively rejects erroneous priors while leveraging correct ones, with accuracy gains that transfer across datasets.