Score
Design and build automated systems and scripts that run web browser engines without a graphical interface to programmatically navigate web pages, control page lifecycle, and inspect or modify the DOM. Implement programmatic interactions (clicks, form input, cookie-banner handling), capture artifacts such as screenshots and DOM snapshots, and collect runtime or compliance metrics at scale.
Web applications exhibit high interface dynamism and complex navigation and form interactions, posing significant challenges for end-to-end test case generation—particularly low coverage and poor robustness. To address this, we propose a navigation modeling approach that integrates screen transition graphs with large language models (LLMs) to enable high-precision path exploration. We further design a state-graph-based automated testing framework for conditional forms, supporting dynamic DOM parsing and interaction modeling. Additionally, we construct the first dedicated benchmark dataset for form-interaction testing. Experimental evaluation across diverse, complex web applications demonstrates substantial improvements in test coverage and path discovery accuracy—especially for branching navigation and conditional form-filling tasks. Our approach establishes a scalable, empirically evaluable paradigm for web application reliability testing.
This work addresses the high learning cost of complex graphical user interfaces and the limitations of existing assistance methods, which often rely on separate chat windows or require extensive custom development. The authors propose an in-situ assistance paradigm that leverages a browser extension to perform lightweight, reversible interventions on the DOM, enabling dynamic interface restructuring without altering the underlying application logic. They introduce the first DOM-based design space and computational pipeline for in-situ assistance, integrating natural language understanding, UI element localization, and reversible operations to support real-time injection of hints, highlighting of controls, or layout rearrangements on arbitrary web pages. Evaluations on two complex interfaces demonstrate the approach’s efficacy and reliability, with user studies showing significant improvements over the ChatGPTAtlas baseline in both usability and task completion efficiency.
Non-programmers face significant challenges in creating secure and efficient software script automations, as conventional approaches require programming expertise and API knowledge, while runtime code generation suffers from unverified outputs, security vulnerabilities, high latency, and substantial computational overhead. Method: This paper proposes an offline simulation-driven framework for skill discovery and validation. It treats software script interfaces as system-level testbeds for large language models (LLMs), employs a graph neural network (GNN)-based API coordination prediction model to identify infrequent yet semantically valid API combinations, and integrates top-down functional guidance with bottom-up API coordination exploration—leveraging offline execution feedback for iterative script refinement. Contribution/Results: Evaluated on Adobe Illustrator, the framework achieves markedly higher automation success rates, significantly reduced response latency, and substantially lower token consumption compared to baseline methods.
This work addresses the vulnerability of existing ReAct-based web agents to prompt injection attacks when encountering malicious web content, which can hijack their control flow. To mitigate this, the authors propose shifting to a “Plan-Then-Execute” paradigm: before observing any web page, the agent generates a task-specific static execution plan and relies on typed, task-level APIs instead of low-level browser operations, thereby isolating untrusted inputs from the decision-making logic. By combining programmatic planning with static task graphs, the approach eliminates runtime dependence on large language models for action generation. Evaluation on the WebArena benchmark shows that all tasks are compatible with this paradigm, and 80% can be completed entirely using pre-generated plans without any runtime LLM invocation, substantially enhancing security and auditability.
This study addresses the limitations of existing web agent evaluation frameworks, which rely on visual scraping or judge models and consequently suffer from inaccurate verification and unquantifiable interface sensitivity. To overcome these issues, this work proposes an open-source, event-log-based evaluation framework that introduces a novel event-level verification mechanism. By leveraging typed log recording and DOM-controlled variant re-rendering, the method achieves precise, judge-free, and scrape-free task assessment across six simulated websites. Furthermore, it supports single-configuration toggling to generate UI variants for quantifying interface sensitivity. The project releases 152 tasks alongside a public leaderboard, revealing a significant discrepancy of up to 41 points between agents' self-reported completion rates and actual performance measured via event logs.
为了解决LLM代理在现有软件界面上效率低下的问题,研究提出了一种名为String的新操作系统,通过将工具知识转化为Markdown文件形式提供给代理使用,从而提高其工作效率。
研究解决了Web代理工具集冗余、与用户需求不匹配的问题,通过AutoTailor框架进行静态筛选和动态重选API,提高了任务准确性并降低了成本。
本文提出RILA,通过执行驱动和浏览器渲染循环,利用AIV模块和ERS评分优化网页代码,提高交互功能和视觉保真度。
论文提出WebWorld接口,通过让VLM与浏览器自主交互并验证代码有效性,解决了VLM驱动的网页代码自我改进中的结构缺陷。
This study addresses the disconnect between existing coding and computer-use agents, as well as the lack of visual interaction to assist software diagnosis and repair, by being the first to systematically investigate the role of visual feedback in this task. Methodologically, it integrates source-code-level execution, application screenshot analysis, and graphical interaction mechanisms to construct a benchmark environment spanning four domains, requiring agents to extract specification information from runtime interfaces and validate their modifications. The primary contribution lies in providing executable correctness evaluation criteria that systematically quantify the capability of state-of-the-art agents to accomplish software engineering tasks by combining code editing, command execution, and GUI-based visual feedback.