🤖 AI Summary
This study addresses the navigation inefficiency of existing mobile GUI agents that rely on screen-by-screen interactions. To handle complex tasks, this work proposes GUI-Hopper, a hybrid interaction architecture that integrates deeplink-based direct navigation with conventional GUI operations. The core innovation lies in a novel method for constructing a deeplink catalog through static code analysis and automated on-device verification, combined with large language model planning to enable hybrid action space modeling. Experimental results demonstrate that this approach significantly improves both task success rates and execution efficiency in real-world commercial applications, validating the effectiveness of the proposed hybrid interaction paradigm.
📝 Abstract
Mobile GUI agents complete tasks using GUI actions like taps and swipes. These actions are broadly applicable across applications, but reaching a navigation interface. A single deeplink call can replace a sequence of screen-by-screen GUI actions. We therefore introduce hybrid interaction, using deeplinks for direct navigation and GUI actions for other on-screen operations and fallback. To enable this, we discover candidate deeplinks through static analysis, validate them on real devices, and describe their observed landing screens. This process creates a verified and grounded deeplink catalog that pairs each working deeplink with a description of its landing screen. Using this catalog, we introduce GUI-Hopper, a improves task success in commercial applications on real devices, further demonstrating the benefits of hybrid interaction.