🤖 AI Summary
Current GUI automation agents suffer from poor generalization, high latency, and weak long-horizon coherence, rendering them fragile to interface changes and inadequate for complex, extended-duration tasks. To address these limitations, we propose an adaptive two-level planning framework: (1) an upper-level, log-driven task mining module that automatically extracts user-specific interaction patterns to construct a structured task dictionary; and (2) a lower-level, vision–semantics joint GUI context understanding module that dynamically maps high-level intentions to executable low-level actions. The framework enables both plan reuse and real-time adaptation, substantially improving robustness and efficiency. Evaluated on 200 real-world tasks, our method achieves a success rate exceeding 60.0% on long-horizon tasks—significantly outperforming state-of-the-art approaches—while simultaneously reducing execution latency.
📝 Abstract
GUI task automation streamlines repetitive tasks, but existing LLM or VLM-based planner-executor agents suffer from brittle generalization, high latency, and limited long-horizon coherence. Their reliance on single-shot reasoning or static plans makes them fragile under UI changes or complex tasks. Log2Plan addresses these limitations by combining a structured two-level planning framework with a task mining approach over user behavior logs, enabling robust and adaptable GUI automation. Log2Plan constructs high-level plans by mapping user commands to a structured task dictionary, enabling consistent and generalizable automation. To support personalization and reuse, it employs a task mining approach from user behavior logs that identifies user-specific patterns. These high-level plans are then grounded into low-level action sequences by interpreting real-time GUI context, ensuring robust execution across varying interfaces. We evaluated Log2Plan on 200 real-world tasks, demonstrating significant improvements in task success rate and execution time. Notably, it maintains over 60.0% success rate even on long-horizon task sequences, highlighting its robustness in complex, multi-step workflows.