🤖 AI Summary
This study addresses the challenges of high visual perception costs and poor execution stability in complex web tasks by proposing a structured-interaction-first web agent system. The method establishes a paradigm that prioritizes structured information over visual perception, integrating semantic web parsing, grid-assisted localization, and hierarchical context management. Furthermore, multi-browser concurrent control and fault-tolerant execution algorithms are designed to enhance robustness. Evaluated on the official Protocol 3 benchmark, the proposed system achieves first place with a 59% pass rate, while attaining a 79% pass rate in local evaluations, thereby demonstrating both the efficiency and stability of the approach.
📝 Abstract
This report presents the web agent system developed for the WebRetriever Challenge. The system follows a structuredinteraction- first strategy, using semantic webpage information for routine browser operations and invoking visual perception only when structured representations are insufficient. Three key designs are introduced: grid-assisted visual localization for difficult-to-access controls, hierarchical context management for reducing redundant page and interaction history, and fault-aware execution mechanisms for stable multi-browser task processing. The system achieved a pass rate of up to 79% in local evaluation on Protocol 1. In the official Protocol 3 competition, it achieved a 59% pass rate with eight concurrent browser workers and ranked first overall, winning the WebRetriever Challenge.