browser instrumentation and automation

Designs and builds tooling and scripts to programmatically control and inspect web browsers using protocols such as the Chrome DevTools Protocol; implements automation for driving page interactions, capturing DOM and network state, collecting runtime and performance diagnostics, and integrating browser-level instrumentation into testing, monitoring, or analysis pipelines.

browserinstrumentationandautomation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the limitations of existing web agent evaluation frameworks, which rely on visual scraping or judge models and consequently suffer from inaccurate verification and unquantifiable interface sensitivity. To overcome these issues, this work proposes an open-source, event-log-based evaluation framework that introduces a novel event-level verification mechanism. By leveraging typed log recording and DOM-controlled variant re-rendering, the method achieves precise, judge-free, and scrape-free task assessment across six simulated websites. Furthermore, it supports single-configuration toggling to generate UI variants for quantifying interface sensitivity. The project releases 152 tasks alongside a public leaderboard, revealing a significant discrepancy of up to 41 points between agents' self-reported completion rates and actual performance measured via event logs.

evaluation benchmarkevent-level verificationtask completion gap

An Empirical Study of JavaScript Inclusion Security Issues in Chrome Extensions

May 26, 2025
CG
Chong Guan
🏛️ Zhejiang Gongshang University

This study systematically uncovers, for the first time, security risks arising from JavaScript (JS) inclusions in Chrome extensions—highlighting critical environmental differences from regular web contexts and a longstanding lack of comprehensive empirical analysis. To address this gap, we propose a hybrid static-dynamic analysis framework: static analysis combines abstract syntax tree (AST) parsing with inter-procedural data-flow path tracing; dynamic analysis integrates runtime script-injection detection and network request monitoring. Applied to 36,324 real-world extensions, our framework identifies 350,784 JS inclusions, confirms 22 exploitable remote script loading vulnerabilities, and reveals widespread reliance on high-risk, outdated libraries (e.g., jQuery 1.x, Underscore 1.4.x) across mainstream extensions. These findings fill a critical gap in browser extension JS supply-chain security research and provide an empirical foundation and methodological support for extension vetting, automated vulnerability detection, and security governance.

Analyzes prevalence of outdated, susceptible JavaScript libraries in popular extensionsIdentifies vulnerable remote JavaScript inclusions enabling arbitrary code executionInvestigates security risks of JavaScript inclusions in Chrome extensions

Current code-generation agents fail to meet functional requirements in over 70% of web application scenarios, primarily due to the absence of automated deployment, browser-level validation, and a feedback loop for iterative repair. This work proposes TDDev, a novel framework that establishes the first fully automated test-driven development (TDD) pipeline by integrating structured acceptance test generation, browser-interaction simulation for validation, and failure-triggered repair signals within a multi-agent collaborative architecture. Empirical evaluation demonstrates that TDDev improves generation quality by 34–48 percentage points, and user studies confirm it eliminates the need for manual intervention entirely. Furthermore, the study reveals that alignment between model generation style and TDD strategy is critical: misalignment not only nullifies performance gains but also inflates token consumption by up to 25-fold.

automated validationbrowser-based testingfunctional correctness

Automated Web Application Testing: End-to-End Test Case Generation with Large Language Models and Screen Transition Graphs

Jun 03, 2025
NL
Nguyen-Khang Le
🏛️ Japan Advanced Institute of Science and Technology | Amifiable Inc.

Web applications exhibit high interface dynamism and complex navigation and form interactions, posing significant challenges for end-to-end test case generation—particularly low coverage and poor robustness. To address this, we propose a navigation modeling approach that integrates screen transition graphs with large language models (LLMs) to enable high-precision path exploration. We further design a state-graph-based automated testing framework for conditional forms, supporting dynamic DOM parsing and interaction modeling. Additionally, we construct the first dedicated benchmark dataset for form-interaction testing. Experimental evaluation across diverse, complex web applications demonstrates substantial improvements in test coverage and path discovery accuracy—especially for branching navigation and conditional form-filling tasks. Our approach establishes a scalable, empirically evaluable paradigm for web application reliability testing.

Automating test case generation for web application navigation using LLMs and screen transition graphsHandling dynamic form interactions in web testing via state graphs and Selenium automationImproving test coverage and robustness in web application reliability assessment

This work investigates whether large language models (LLMs) can synthesize executable, reusable Selenium Python scripts from natural-language goals by parsing HTML/DOM structures to enable end-to-end web automation. To this end, we introduce the first code-first, task-driven benchmark for web automation code generation—covering 7 major website categories and 681 multi-step tasks—and propose an end-to-end verification protocol integrating static analysis, sandboxed execution, result assertion, and security safeguards. Experimental evaluation shows that GPT-4o-Mini achieves a functional success rate of 96.8% across 2,636 runs (91.7% on average for simple tasks), yet fails entirely on complex workflows; critically, no model produces production-grade code. This study constitutes the first systematic assessment of LLMs’ program synthesis capability in realistic web environments, establishing a standardized testbed and empirically grounded benchmark for browser automation and AI agent research.

Evaluating LLMs' ability to synthesize reusable browser automation programsTesting web automation scripts across 681 tasks on self-hosted sitesValidating generated code through execution and safety verification

Latest Papers

What's happening recently
View more

This work addresses the challenge of automatically reproducing web GUI bug reports, which often lack critical contextual information such as dependencies or input files. To this end, the authors propose ReBug, the first end-to-end, state-aware bug reproduction agent for web GUIs. ReBug operates in two stages: it first leverages a large language model to reconstruct missing context and generate a high-level reproduction plan, then executes state-aware interactions within a real browser, validating outcomes through structured page summaries and a history replay mechanism. Evaluated on 667 real-world bugs, ReBug achieves an average Reproduction Success Rate (RSR) of 49.96% and a Task Completion Rate of 74.96%, significantly outperforming existing approaches.

browser executionbug reportsprerequisites reconstruction

为了解决LLM代理在现有软件界面上效率低下的问题,研究提出了一种名为String的新操作系统,通过将工具知识转化为Markdown文件形式提供给代理使用,从而提高其工作效率。

efficiencyLLM agentsresource waste

Hot Scholars

JC

Junyu Cao

McCombs School of Business, The University of Texas at Austin
JN

Jingyi Ni

Software Engineer, Airbnb
Embodied AIAutonomous Agent
KC

Kunyu Chen

Engineering Manager (Ads ML Infra), Meta Platforms
DZ

Danqing Zhang

University of California, Berkeley
Natural Language ProcessingCyber Physical SystemsAutonomous AgentsEmbodied AI