Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether large language models (LLMs) tend to apply localized patches or complete rewrites when repairing code, and how their behavior diverges from that of human developers. To this end, the authors construct an evaluation benchmark based on Codeforces and present the first quantitative analysis measuring the deviation of LLM repair behaviors from human-generated patches. Furthermore, they conduct automated comparative evaluations across GPT-series models using text similarity algorithms. The findings reveal that LLMs inherently violate the principle of minimal modification, exhibiting a strong preference for reimplementation over targeted patching. Notably, the study demonstrates that generating solutions from scratch outperforms repairing erroneous code. These insights provide critical implications for the design of AI-assisted debugging tools, highlighting fundamental discrepancies between machine-driven and human-driven code repair strategies.
📝 Abstract
Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in problem solving or bug fixing independently, but do not explore the relationship between these two capabilities. This work focuses on determining how much the LLM deviates from a buggy solution to fix the bug compared to a human-written patch, and if there is a bias towards generating entirely new solutions. We construct a dataset with all the submissions ($\sim$ 3000) from a couple of users from Codeforces, and we match each buggy submission with its corresponding human fix. By using the similarity between the buggy solution and the human fix as a baseline, we evaluate the quality of LLM-generated bug fixes on 3 OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1). We check if the generated solutions solve the problem by using the Codeforces-R1 dataset, an openly available dataset that has tests generated with the DeepSeek-R1 model. Our findings suggest that LLMs tend to modify more lines than necessary compared to human fixes and, in some cases, generate entirely new solutions. We also observe that LLMs solve more problems correctly when allowed to generate solutions from scratch rather than patch buggy submissions, even when those submissions are close to the human patch. This has important implications for the design of AI-assisted programming tools, particularly in supporting user debugging processes and promoting incremental problem-solving strategies rather than solution replacement.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Bug Fixing
Code Repair
Programming
Solution Reimplementation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Bug Fixing
Code Similarity
Competitive Programming
AI-Assisted Debugging
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alexandru Stefan Stoica
University Politehnica of Bucharest, Bucharest, Romania
Traian Rebedea
Traian Rebedea
NVIDIA & Assoc Prof @ University Politehnica of Bucharest
Artificial IntelligenceNatural Language ProcessingMachine LearningHuman-Computer Interaction
M
Marian Cristian Mihaescu
University of Craiova, Craiova, Romania