How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks

📅 2026-04-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of current code generation evaluations, which typically rely on single-attempt generation and overlook the potential of large language models to iteratively correct errors. The authors systematically assess the self-repair capabilities of seven state-of-the-art models—spanning architectures from 8B parameters to MoE with 128 experts—on the HumanEval and MBPP benchmarks, allowing up to five attempts with execution feedback. They demonstrate for the first time that modern instruction-tuned models can achieve effective self-repair through prompting alone, reveal how error types influence repair difficulty, and quantify the performance gains from chain-of-thought prompting. Results show consistent improvements across all models: pass@5 scores increase by 4.9–17.1 percentage points on HumanEval and 16.0–30.0 on MBPP, with Gemini 2.5 Flash achieving final pass rates of 96.3% and 93.8%, respectively.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageComputer Vision: Large Vision Models

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Large language models frequently fail to produce correct code on their first attempt, yet most benchmarks evaluate them in a single-shot setting. We investigate iterative self-repair (feeding execution errors back to the model for correction) across seven models spanning three families and both open-weight and proprietary providers: Llama 3.1 8B, Llama 3.3 70B, Llama 4 Scout (MoE, 16 experts), Llama 4 Maverick (MoE, 128 experts), Qwen3 32B, Gemini 2.5 Flash, and Gemini 2.5 Pro. On HumanEval (164 problems) and MBPP Sanitized (257 problems) with up to five attempts, self-repair universally improves pass rates: +4.9 to +17.1 pp on HumanEval and +16.0 to +30.0 pp on MBPP. Gemini 2.5 Flash achieves the highest final pass rates (96.3% HumanEval, 93.8% MBPP). Most gains concentrate in the first two rounds.Error-type analysis shows assertion errors (logical mistakes) are the hardest to repair at ~45%, while syntax and name errors are repaired at substantially higher rates, connecting to broader findings on the limits of LLM self-correction. Prior work found that weaker models fail at self-repair or require fine-tuning; we show that modern instruction-tuned models succeed with prompting alone, even at 8B scale. We also provide the first comparison of dense and MoE architectures for self-repair, and extend the repair-vs-resampling tradeoff analysis to modern models. A prompt ablation reveals chain-of-thought repair yields up to +5.5 pp additional self-repair gain (measured as improvement in repair delta) over minimal prompting for capable models.
Problem

Research questions and friction points this paper is trying to address.

iterative self-repair
code generation
large language models
benchmark evaluation
execution errors
Innovation

Methods, ideas, or system contributions that make the work stand out.

iterative self-repair
code generation
mixture-of-experts
chain-of-thought prompting
error-type analysis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Johin Johny Arimbur
Independent Researcher