Score
Designs and builds tooling and scripts to programmatically control and inspect web browsers using protocols such as the Chrome DevTools Protocol; implements automation for driving page interactions, capturing DOM and network state, collecting runtime and performance diagnostics, and integrating browser-level instrumentation into testing, monitoring, or analysis pipelines.
This study addresses the limitations of existing web agent evaluation frameworks, which rely on visual scraping or judge models and consequently suffer from inaccurate verification and unquantifiable interface sensitivity. To overcome these issues, this work proposes an open-source, event-log-based evaluation framework that introduces a novel event-level verification mechanism. By leveraging typed log recording and DOM-controlled variant re-rendering, the method achieves precise, judge-free, and scrape-free task assessment across six simulated websites. Furthermore, it supports single-configuration toggling to generate UI variants for quantifying interface sensitivity. The project releases 152 tasks alongside a public leaderboard, revealing a significant discrepancy of up to 41 points between agents' self-reported completion rates and actual performance measured via event logs.
This study systematically uncovers, for the first time, security risks arising from JavaScript (JS) inclusions in Chrome extensions—highlighting critical environmental differences from regular web contexts and a longstanding lack of comprehensive empirical analysis. To address this gap, we propose a hybrid static-dynamic analysis framework: static analysis combines abstract syntax tree (AST) parsing with inter-procedural data-flow path tracing; dynamic analysis integrates runtime script-injection detection and network request monitoring. Applied to 36,324 real-world extensions, our framework identifies 350,784 JS inclusions, confirms 22 exploitable remote script loading vulnerabilities, and reveals widespread reliance on high-risk, outdated libraries (e.g., jQuery 1.x, Underscore 1.4.x) across mainstream extensions. These findings fill a critical gap in browser extension JS supply-chain security research and provide an empirical foundation and methodological support for extension vetting, automated vulnerability detection, and security governance.
Current code-generation agents fail to meet functional requirements in over 70% of web application scenarios, primarily due to the absence of automated deployment, browser-level validation, and a feedback loop for iterative repair. This work proposes TDDev, a novel framework that establishes the first fully automated test-driven development (TDD) pipeline by integrating structured acceptance test generation, browser-interaction simulation for validation, and failure-triggered repair signals within a multi-agent collaborative architecture. Empirical evaluation demonstrates that TDDev improves generation quality by 34–48 percentage points, and user studies confirm it eliminates the need for manual intervention entirely. Furthermore, the study reveals that alignment between model generation style and TDD strategy is critical: misalignment not only nullifies performance gains but also inflates token consumption by up to 25-fold.
Web applications exhibit high interface dynamism and complex navigation and form interactions, posing significant challenges for end-to-end test case generation—particularly low coverage and poor robustness. To address this, we propose a navigation modeling approach that integrates screen transition graphs with large language models (LLMs) to enable high-precision path exploration. We further design a state-graph-based automated testing framework for conditional forms, supporting dynamic DOM parsing and interaction modeling. Additionally, we construct the first dedicated benchmark dataset for form-interaction testing. Experimental evaluation across diverse, complex web applications demonstrates substantial improvements in test coverage and path discovery accuracy—especially for branching navigation and conditional form-filling tasks. Our approach establishes a scalable, empirically evaluable paradigm for web application reliability testing.
This work investigates whether large language models (LLMs) can synthesize executable, reusable Selenium Python scripts from natural-language goals by parsing HTML/DOM structures to enable end-to-end web automation. To this end, we introduce the first code-first, task-driven benchmark for web automation code generation—covering 7 major website categories and 681 multi-step tasks—and propose an end-to-end verification protocol integrating static analysis, sandboxed execution, result assertion, and security safeguards. Experimental evaluation shows that GPT-4o-Mini achieves a functional success rate of 96.8% across 2,636 runs (91.7% on average for simple tasks), yet fails entirely on complex workflows; critically, no model produces production-grade code. This study constitutes the first systematic assessment of LLMs’ program synthesis capability in realistic web environments, establishing a standardized testbed and empirically grounded benchmark for browser automation and AI agent research.
研究解决了Web代理工具集冗余、与用户需求不匹配的问题,通过AutoTailor框架进行静态筛选和动态重选API,提高了任务准确性并降低了成本。
本文使用Cypress框架对开源Web应用进行端到端自动化测试评估,通过27个测试案例分析了执行速度、可靠性和可维护性,证明了Cypress的有效性。
本文针对现代Web应用的动态性和交互性,提出SpiderSapien,一种客户端中心的爬虫和安全扫描器,通过沉浸式交互提高代码覆盖率和漏洞检测率。
This work addresses the challenge of automatically reproducing web GUI bug reports, which often lack critical contextual information such as dependencies or input files. To this end, the authors propose ReBug, the first end-to-end, state-aware bug reproduction agent for web GUIs. ReBug operates in two stages: it first leverages a large language model to reconstruct missing context and generate a high-level reproduction plan, then executes state-aware interactions within a real browser, validating outcomes through structured page summaries and a history replay mechanism. Evaluated on 667 real-world bugs, ReBug achieves an average Reproduction Success Rate (RSR) of 49.96% and a Task Completion Rate of 74.96%, significantly outperforming existing approaches.
为了解决LLM代理在现有软件界面上效率低下的问题,研究提出了一种名为String的新操作系统,通过将工具知识转化为Markdown文件形式提供给代理使用,从而提高其工作效率。