Same Tasks, Different Apps: Why Mobile GUI Agents Fail to Generalize?

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited cross-application generalization of mobile GUI agents and the inadequacy of existing benchmarks in evaluating this deficiency. To this end, it proposes AnyAppBench, a real-time Android benchmark spanning 52 applications that assesses agent performance across heterogeneous interfaces through category control and fixed-target testing. By integrating VLM-as-a-judge evaluation, automated task template generation, and human-in-the-loop annotation, the work establishes a systematic failure taxonomy. The contributions include quantitatively revealing cross-application generalization gaps among thirteen agents, demonstrating that success on source applications is difficult to transfer, and showing that subgoal decomposition strategies yield limited effectiveness. These findings provide critical analytical foundations for improving the robust deployment of mobile GUI agents.
📝 Abstract
Mobile GUI agents deployed in real settings must work across different applications that support the same functionality. Most existing benchmarks test each task in only one app, so a high score can mean the agent understands the task, or only that it knows that particular app. We introduce AnyAppBench, a category-controlled live Android benchmark that evaluates cross-application generalization while keeping the user goal fixed. It spans 10 functional categories, 100 task templates, and 520 task--application pairs over 52 applications. Agents run from raw instructions and with app-independent sub-goals, and a VLM judge labels every failed run under a fixed failure taxonomy whose reliability is measured by human annotation. We find that, across 13 agents, success on the original application does not transfer reliably to new applications with the same goal. Furthermore, providing high-level sub-goal decomposition produces only small, category-dependent changes that do not close the gap, and the mix of failure types changes with the target interface. Based on those insights, we believe the AnyAppBench benchmark provides an important stepping stone toward robust real-world deployment of mobile GUI agents. Our code, data and the leaderboard can be found at the project website https://anyappbench.github.io/.
Problem

Research questions and friction points this paper is trying to address.

Mobile GUI Agents
Cross-application Generalization
Benchmark Evaluation
Task Transferability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mobile GUI Agents
Cross-application Generalization
Benchmark
VLM Judge
Failure Taxonomy
💼 Related Jobs
No related jobs found.
T
Tien Tran
KAIST
N
Namho Koh
KAIST
D
Daiki E. Matsunaga
KAIST
A
Ayush Jain
CMU
K
Kee Eung Kim
KAIST