Towards the Generalizability of Leveraging ChatGPT in APR via Self-enhancing: An Empirical Study

📅 2026-09-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了ChatGPT增强的自动程序修复方法在不同基准测试中的稳定性问题,通过对比实验和代码转换等手段,发现直接提供错误信息可能比现有增强方法更有效。
📝 Abstract
Automated Program Repair (APR) increasingly relies on Large Language Models (LLMs). ChatGPT-enhanced APR uses techniques such as self-correction and autonomous agents to improve repair without modifying model parameters. Although these approaches report strong results on Defects4J and SWE-bench, the stability of enhancement gains across benchmarks remains under-explored. We evaluate three ChatGPT-enhanced APR methods on three representative, long-standing benchmarks. With GPT-3.5-Turbo, SRepair achieves a larger absolute gain on HumanEval-Java than on Defects4J, while SRepair and FixAgent without extrinsic information yield negative gains on BugsInPy. With GPT-5.4-mini, the evaluated methods achieve larger absolute gains on Defects4J than on HumanEval-Java, while gains on BugsInPy are non-negative but limited. We investigate benchmark-related factors through code transformations and benchmark-specific fine-tuning. Code transformations reduce enhancement gains on Defects4J, while benchmark-specific fine-tuning increases gains on BugsInPy. Directly supplying GPT-3.5-Turbo with error messages and triggering tests yields more correct repairs than the evaluated ChatGPT-enhanced APR methods on BugsInPy. These findings highlight the need to evaluate generalizability across benchmarks and models using multiple metrics, and suggest that directly providing repair-specific extrinsic information may be more effective than enhancement methods when their gains are limited.
Problem

Research questions and friction points this paper is trying to address.

Automated Program Repair
Large Language Models
Generalizability
Benchmarks
Stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-enhancing
Automated Program Repair (APR)
Large Language Models (LLMs)
benchmark-specific fine-tuning
repair-specific extrinsic information
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Qingyuan Li
Qingyuan Li
Meituan
AutoMLNeural Network CompressionHardware AccelerationLarge Language ModelAIGC
C
Chuanyi Li
State Key Laboratory for Novel Software Technology, Nanjing University, Jiangsu, China
Y
Yaopeng Yang
State Key Laboratory for Novel Software Technology, Nanjing University, Jiangsu, China
Z
Ziwen Ge
State Key Laboratory for Novel Software Technology, Nanjing University, Jiangsu, China
Jidong Ge
Jidong Ge
Associate Professor of Software Engieering, Nanjing University
Software EngieeringWorkflowServices ComputingPetri Nets
Bin Luo
Bin Luo
Assistant Professor, Southern University of Science and Technology
earthquake physicsgeomechanicsseismologydistributed acoustic sensing