🤖 AI Summary
This study addresses the vulnerability of GUI agents to attention diversion and cascading failures caused by abrupt interruptions in long-horizon dynamic scenarios. To this end, it proposes a pre-cognitive architecture that shifts the paradigm from passive reaction to proactive decision-making. Methodologically, the approach introduces a dual-memory experience pool to cache anomalous patterns, constructs a symbolic UI simulation executor to predict interface evolution, and employs a controller that integrates prior knowledge with predictive signals to guide robust decision-making. Additionally, an AutoTraj data generation engine and InterfereBench, a benchmark featuring strong interruptions, are presented. Experimental results demonstrate that the proposed method significantly outperforms existing state-of-the-art approaches on InterfereBench while maintaining competitive performance on public benchmarks, thereby substantially enhancing robustness in long-horizon tasks.
📝 Abstract
Existing reactive Graphical User Interface (GUI) agents often fail in long-horizon, dynamic scenarios, where unexpected disturbances trigger attention-diverting and cascading failures. To address this, we propose PrecogUI, a pre-cognitive architecture that shifts the paradigm from reactive execution to proactive decision-making. Specifically, we design a Proactive Experience Pool (PEP), which caches recurring anomaly and success patterns as "state-action-result" tuples in a dual-memory repository. Furthermore, we introduce a Proactive Simulation Executor (PSE) that learns to forecast the next symbolic UI layout given a candidate action, enabling early anomaly avoidance and ranking candidate actions by predicted reliability. Finally, a Pre-cognitive Execution Controller (PEC) fuses these priors and predictions, prioritizes handling of foreseen anomalies, and ensures execution robustness through a closed-loop error correction mechanism. For robust evaluation, we develop AutoTraj, an automatic data-generation engine, to construct InterfereBench, a benchmark for long-horizon tasks with strong disturbances. Experiments demonstrate that PrecogUI surpasses state-of-the-art methods on InterfereBench while maintaining competitive performance on public benchmarks. The code will be publicly available.