Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance on manual verification for codeless bug fixes generated by large language models, which is both time-consuming and error-prone. To overcome this limitation, we propose an end-to-end automated evaluation framework grounded in real browser execution. The approach leverages a computer-use agent (CUA) to automatically execute natural language repair instructions and verify whether the reported issues are resolved, while systematically quantifying how different agents influence fix success rates. Experimental results demonstrate that the optimal configuration achieves a 74.1% repair rate, whereas the overall average remains below 50%. Furthermore, variations in the executor significantly affect evaluation outcomes, confirming both the necessity and practical value of automated, execution-based verification.
📝 Abstract
A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow. Manually verifying whether a proposed no-code fix resolves the reported bug takes considerable developer time. This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment. We evaluate 322 no-code fixes generated by the 12 configurations released with the benchmark of a previous study, covering bug reports categorized as Faulty Configuration, Wrong Version, or External System & Dependency. An executor agent applies each fix by following its natural-language instructions, and an issue-specific checker determines whether the reported bug persists. We repeat the pipeline with three executors: two Computer-Use Agents, OpenCUA-72B and Claude Sonnet 5, and one multimodal agentic LLM, Meta's Muse Glimmer. Only 17.6% of the candidate issues could be set up and passed both sanity gates. Across the 322 fixes, 14.6% to 49.7% resolved the bug depending on the executor, and the strongest configuration, Claude Opus 4.6 in the Vanilla pipeline, resolved up to 74.1% of its fixes under Claude Sonnet 5. Changing only the executor shifted a configuration's resolution rate by 38.8% on average, and the three executors reached the same verdict on only 46.9% of the fixes. Compared with human execution, the executors matched the human consensus for 66.1% to 88.1% of the sampled fixes. Even under the best executor, fewer than half of the LLM-generated no-code fixes resolve the reported bug, so such fixes need verification before they reach users. Execution-based verification can provide this, but the measured capability depends strongly on the executor, which evaluations must report and control.
Innovation

Methods, ideas, or system contributions that make the work stand out.

No-Code Bug Fixes
Automated Verification
Execution-based Pipeline
Computer-Use Agents
Large Language Models