🤖 AI Summary
This study addresses the insufficient cross-user reliability of existing mobile GUI agents operating within personalized interfaces. To bridge this gap, we propose the PAIR evaluation pipeline and the RePAIR training framework. Specifically, we introduce the first user-state instantiation pipeline to quantify performance disparities across users, and integrate application-state rendering, subgoal reward modeling, and reinforcement learning fine-tuning to achieve personalization-aware optimization. Experimental results demonstrate that our approach effectively mitigates cross-user discrepancies. On unseen-user test sets, it yields improvements of 5.87, 7.50, and 9.42 percentage points in step-level action match rate (SAR), full success rate, and overall task success rate, respectively.
📝 Abstract
Mobile GUI agents increasingly operate on interfaces influenced by users' histories and preferences, but their reliability across different users remains underexplored. We introduce PAIR (Personalized Application-state Instantiation and Rendering), a pipeline for constructing user-conditioned application states that enables controlled evaluation of the same task across different users. We further introduce RePAIR (Reinforcement learning with Personalization-Aware Interaction Rewards), a training approach that learns from cross-user differences in subgoal outcomes to improve reliability across user-conditioned mobile environments. Across six agents, we find substantial variation in task success across users and consistently lower subgoal achievement in user-conditioned UI contexts (6.98 to 15.4 pp). This gap further increases for personal targets drawn from each user's own content (8.77 to 22.0 pp). Failures in these contexts frequently involve selecting another item instead of the intended target, particularly before target exposure. Finally, RePAIR improves user-conditioned SAR (+5.87 pp), all-success (+7.50 pp), and overall Task SR (+9.42 pp) over its supervised fine-tuning parent on unseen users, providing initial evidence that explicitly learning from cross-user variation can improve GUI-agent reliability.