Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?
This study investigates whether large language models (LLMs) tend to apply localized patches or complete rewrites when repairing code, and how their behavior diverges from that of human developers. To this end, the authors construct an evaluation benchmark based on Codeforces and present the first quantitative analysis measuring the deviation of LLM repair behaviors from human-generated patches. Furthermore, they conduct automated comparative evaluations across GPT-series models using text similarity algorithms. The findings reveal that LLMs inherently violate the principle of minimal modification, exhibiting a strong preference for reimplementation over targeted patching. Notably, the study demonstrates that generating solutions from scratch outperforms repairing erroneous code. These insights provide critical implications for the design of AI-assisted debugging tools, highlighting fundamental discrepancies between machine-driven and human-driven code repair strategies.