Diff-Based Code Corruption using LLMs for Large-Scale Bugfix Benchmarking

📅 2026-06-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing code repair benchmarks suffer from limited scale, insufficient diversity, and a narrow range of bug types, making it difficult to accurately assess the repair capabilities of large language models. To address this, this work proposes a diff-based, controllable code corruption method that leverages large language models to synthesize semantically plausible buggy programs from correct ones, thereby avoiding overly simplistic or invalid bugs. Using this approach, we construct MegaBugFix, a large-scale, highly diverse benchmark comprising 12,629 Python defect samples. Evaluation of 13 open-source models on MegaBugFix reveals substantially lower performance compared to existing benchmarks, highlighting the limitations of current models in handling complex, realistic bugs and underscoring the challenge and evaluation value of our benchmark.
📝 Abstract
There are various benchmarks to evaluate bugfixing capabilities of Large Language Models. However, most widespread benchmarks do not fully reflect real-world bugfixing practices. They are small, weakening statistical reliability, and the buggy programs are often similar to one another, potentially distorting evaluation results. The range of bug types can also be narrow, failing to capture a representative range of bugs. To address these issues, we introduce MegaBugFix, a large-scale bugfixing benchmark containing 12,629 buggy Python programs synthesized from correct ones by a Large Language Model. Bug injections were generated as diffs representing code changes. Through this approach, we were able to avoid common pitfalls of LLM-based mutation techniques like injecting overly simplistic bugs or failing to modify the input program. We evaluated 13 open-weight models on MegaBugFix and baseline benchmarks, finding consistently lower performance on MegaBugFix. This reveals that our benchmark presents more challenging bugs and exposes model failures that may remain hidden when evaluating on existing benchmarks. The benchmark and fine-tuned model used for bug injection are available at hf.co/collections/szalontaib/megabugfix.
Problem

Research questions and friction points this paper is trying to address.

bugfixing benchmark
Large Language Models
code corruption
real-world bugfixing
benchmark limitations
Innovation

Methods, ideas, or system contributions that make the work stand out.

diff-based bug injection
large-scale benchmarking
LLM-generated bugs
code mutation
MegaBugFix
💼 Related Jobs
No related jobs found.
B
Balázs Szalontai
Eötvös Loránd University, Faculty of Informatics
Á
Ábel Szauter
Eötvös Loránd University, Faculty of Informatics
B
Balázs Márton
Eötvös Loránd University, Faculty of Informatics
P
Péter Verebics
Eötvös Loránd University, Faculty of Informatics
Balázs Pintér
Balázs Pintér
Eötvös Loránd University
Machine LearningNatural Language Processing
T
Tibor Gregorics
Eötvös Loránd University, Faculty of Informatics