From Bug Reports to Browser-Executable Procedures: An LLM-Driven Agent for Web GUI Bug Reproduction

📅 2026-08-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of automatically reproducing web GUI bug reports, which often lack critical contextual information such as dependencies or input files. To this end, the authors propose ReBug, the first end-to-end, state-aware bug reproduction agent for web GUIs. ReBug operates in two stages: it first leverages a large language model to reconstruct missing context and generate a high-level reproduction plan, then executes state-aware interactions within a real browser, validating outcomes through structured page summaries and a history replay mechanism. Evaluated on 667 real-world bugs, ReBug achieves an average Reproduction Success Rate (RSR) of 49.96% and a Task Completion Rate of 74.96%, significantly outperforming existing approaches.
📝 Abstract
Reproducing web GUI bugs from natural-language bug reports is critical for software maintenance, but remains difficult because reports often lack prerequisites such as dependencies and input files. Existing bug reproduction techniques mainly target code units or mobile applications and lack end-to-end visual execution and validation for web GUIs. We present ReBug, a context-aware agent system that reconstructs, executes, and validates browser-level reproduction procedures from web GUI bug reports by driving a real browser. ReBug separates reproduction into two stages. In the preparation stage, ReBug reconstructs missing prerequisites from the report and available artifacts, and it produces a high-level reproduction plan. In the execution stage, it performs tool-mediated interactions in the browser, maintains structured summaries of page state and action history, and validates the final state against expectations derived from the report. We evaluate ReBug on 667 real-world bug reports from four open-source web applications. On controlled current deployments, ReBug outperforms both baselines, achieving an average RSR of 49.96%, a mean task completion rate of 74.96%, and a mean action execution success rate of 86.54%. Our results show that explicit context reconstruction and state-aware browser execution effectively support report-derived browser reproduction, while historical replay shows that successful procedures often expose the original bug-present behavior on restored buggy versions.
Problem

Research questions and friction points this paper is trying to address.

Web GUI bug reproduction
bug reports
browser execution
software maintenance
prerequisites reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-driven agent
web GUI bug reproduction
context reconstruction
browser-level execution
state-aware validation
C
Cunming Zhang
University of Luxembourg
Y
Yu Pei
University of Luxembourg
M
Michail Papadakis
University of Luxembourg