🤖 AI Summary
This study addresses two limitations in reinforcement learning for GUI agents: binary evaluation overlooks spatial discrepancies in click predictions, and entirely failed trajectories lack relative supervisory signals. To overcome these issues, this work proposes a spatial credit assignment method integrated within a group relative reinforcement learning framework. By incorporating screen coordinate regression prediction residuals and a distance-based ranking mechanism, the approach refines intra-group relative credits through spatial references, effectively mitigating sparse rewards and vanishing gradient signals while optimizing policy updates. Experimental results demonstrate that the proposed method significantly improves grounding accuracy across multiple benchmarks, achieving state-of-the-art action prediction performance and delivering consistent gains across diverse domains.
📝 Abstract
GUI agents automate tasks on digital devices by grounding language instructions in visual interfaces. Existing group-relative reinforcement learning improves GUI action prediction by comparing the rewards of multiple responses sampled from the same GUI state. However, binary evaluation treats spatially different failed clicks as identical and provides no relative signal when all sampled clicks fail. To address these limitations, we propose Spatial Credit Assignment (SCA), which uses the screen coordinates of sampled clicks to refine group-relative credit. Specifically, SCA predicts each held-out response's reward from the other responses in groups containing both successes and failures, then uses the prediction residual to adjust credit. When all sampled clicks fail, SCA instead orders them by distance to the annotated target. These spatial references are used only to construct the training update; the deployed policy remains unchanged. We evaluate whether this correction improves the policy update itself by comparing its error and directional alignment with the exact return gradient in a controlled synthetic study. Across GUI grounding and offline action-prediction benchmarks, SCA improves grounding across professional domains and achieves the strongest results among reinforcement-fine-tuned models on most action-prediction metrics, with consistent gains across the reported GUI suites.