Merged, Not Measured: An Empirical Study of Performance Issues Fixed by Coding Agents

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the practical effectiveness of coding agents in resolving performance issues and the mechanisms governing maintainer acceptance. By analyzing over 70,000 agent-generated pull requests and identifying 1,262 performance-related fixes, the research employs a combination of text filtering, LLM-assisted classification, and empirical reproduction testing. Results indicate that while 57% of these fixes are merged, maintainer decisions rely predominantly on repository history rather than code quality. Furthermore, verification reveals that most merged patches fail to achieve the anticipated performance improvements. The analysis identifies redundant computation as the primary source of performance defects and demonstrates that such fixes frequently introduce substantial behavioral change risks. These findings provide critical empirical evidence for guiding agent-based performance repair practices in software maintenance.
📝 Abstract
Coding agents open pull requests (PRs) that claim to speed up software, but studies of human performance fixes say little about how maintainers respond to such a fix or whether its claim holds. From the 71,677 agent PRs of AIDev v4, a text filter and codebook coding by language models and by the authors select 1,262 performance issues fixed by six agents in 582 repositories. We code each issue and its tests and re-execute 23 rejected and 30 merged fixes. (1) 57% of closed fixes are merged, 61% of rejections give no stated reason, and only 6 of the 23 re-executed rejected claims held under our three-run pilot on mostly agent-built workloads. (2) Acceptance rises with the agent's track record in the repository (31-37% to 70%) and with the repository's pre-opening merge rate on its other agent PRs (33% to 84%). Merged fixes delete a larger share of the lines they change (0.26 versus 0.15), a difference that holds within agent and within repository, with no such difference detected in the coded content, description, tests or measurements. (3) Repeated computation and redundant data processing cause 44% of the issues, and 46% of fixes are architectural-level. (4) Agents change tests in 37% of fixes and 11% carry a performance test or benchmark; of the 30 merged fixes, 18 met our delivery criterion, 3 fell short of the claim, 9 showed no significant gain or regressed, and 14 change behavior on untested inputs. The outcome tracks the repository's history with the agent rather than the coded content of the fix, and a merge does not show that the fix delivers what it claims.
Problem

Research questions and friction points this paper is trying to address.

coding agents
performance issues
pull requests
empirical study
software maintenance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Coding Agents
Performance Issues
Empirical Study
Pull Requests
Automated Program Repair
🔎 Similar Papers