Score
Designs, builds, and evaluates automated repair loops for models (including large language models and other probabilistic models) that use confidence or calibration metrics to prioritize, generate, and apply targeted edits to outputs, likelihoods, priors, or parameterizations. This competence covers implementing calibration-guided or LLM-driven iterative repair procedures, and analysing repair outcomes and functional regressions to ensure improved calibration and reliable behavior after fixes.
This work addresses the high cost, poor scalability, and diminishing effectiveness of human-supervised approaches for improving large language models, especially as model capabilities approach human-level performance. To overcome these limitations, the paper proposes a closed-loop self-improvement framework that structures the self-enhancement process into four tightly coupled stages: data acquisition, selection, model optimization, and inference refinement. A key innovation is the introduction of an autonomous evaluation layer that coordinates and guides transitions across these stages. This framework offers the first systematic, lifecycle-oriented modeling of self-improvement, unifying critical components such as self-generated data, automated evaluation, iterative training, and inference-time optimization. By comprehensively mapping existing technical pathways and their limitations, the study lays the groundwork for realizing fully autonomous, self-evolving language models.
This paper addresses the challenge of automated repair of C programs by proposing the first closed-loop framework that integrates spectrum-based fault localization (SBFL) with large language model (LLM)-driven chain-of-thought (CoT) reasoning. Methodologically, the approach employs runtime test feedback to drive iterative refinement: SBFL initially identifies suspicious statements; an LLM generates patches using structured CoT prompting; and execution feedback guides successive improvements until all tests pass. The key contribution lies in the explicit coupling of statistical program analysis with symbolic LLM reasoning—enhancing fault attribution accuracy and mitigating recurrent patch failures. Evaluated on the Codeflaws benchmark (3,902 defects), the method achieves a repair accuracy of 44.93%, outperforming the GPT-4+CoT baseline by 3.61%, thereby validating the efficacy of analysis-guided iterative reasoning for program repair.
This work addresses the challenge that probabilistic programs generated by language models often suffer from statistical misspecifications—such as incorrect likelihoods, priors, or parameterizations—that are difficult to detect with conventional unit tests. The paper introduces, for the first time, Bayesian calibration as a central criterion for assessing the correctness of probabilistic programs and proposes a fully unsupervised, reference-free framework for their detection and repair. By integrating Bayesian validation techniques—including posterior predictive checks, simulation-based calibration (SBC), sampling diagnostics (e.g., $\hat{R}$, divergences, effective sample size), and held-out predictive log density—the method generates feedback signals to drive an iterative repair loop within large language models. Evaluated on 200 instances, the approach achieves detection AUCs of 0.97 with reference programs and 62–78% without, substantially outperforming unit testing; repair success rates reach 92% and 100% using GPT-5.1 and Claude, respectively.
This study systematically evaluates the automatic program repair (APR) capabilities of four open-source large language models—CodeLlama, LLaMA, StarCoder, and DeepSeek-Coder—across programming languages (Java, C, C++, Python) and defect types (industrial vs. algorithmic). Using six benchmark datasets, four prompting strategies, and multi-dimensional patch evaluation, we conduct a large-scale empirical analysis over 600,000 generated patches. Our findings reveal: (1) code-specialized models significantly outperform general-purpose models of comparable scale; (2) repair performance exhibits nonlinear scaling with model size, and optimal patches are disproportionately concentrated in early-generation outputs; and (3) prompt engineering improves repair rates by over 40%. This work advances beyond traditional APR evaluations—typically limited to small models and single-language settings—by uncovering critical principles governing model specialization, prompt sensitivity, and solution-space distribution in large-language-model-based APR.
This work presents the first systematic evaluation of large language models’ (LLMs) ability to infer and repair program loop invariants without auxiliary information. We adopt an empirical framework encompassing diverse open- and closed-source LLMs across multiple scales, integrating domain-knowledge augmentation and few-shot prompting to quantify performance on standard benchmarks for inductive invariant generation and logical defect repair. Results show that LLMs achieve up to 78% success in invariant generation but only 16% in invariant repair—revealing a critical bottleneck in deep logical correction. A key contribution is the identification of auxiliary information—particularly loop semantics prompts and correct examples—as decisive for improving repair accuracy. Our study establishes a reproducible evaluation paradigm for LLM-driven automated program safety analysis and provides concrete, actionable pathways for enhancing invariant repair capabilities.
This study addresses the limitation of current code generation evaluations, which typically rely on single-attempt generation and overlook the potential of large language models to iteratively correct errors. The authors systematically assess the self-repair capabilities of seven state-of-the-art models—spanning architectures from 8B parameters to MoE with 128 experts—on the HumanEval and MBPP benchmarks, allowing up to five attempts with execution feedback. They demonstrate for the first time that modern instruction-tuned models can achieve effective self-repair through prompting alone, reveal how error types influence repair difficulty, and quantify the performance gains from chain-of-thought prompting. Results show consistent improvements across all models: pass@5 scores increase by 4.9–17.1 percentage points on HumanEval and 16.0–30.0 on MBPP, with Gemini 2.5 Flash achieving final pass rates of 96.3% and 93.8%, respectively.
This study addresses the lack of systematic understanding regarding the impact of repair loop iteration counts in large language model (LLM)-based software engineering tasks, where prior work often relies on arbitrarily defined repair budgets. Through a cross-task (code generation, test generation, code translation) and cross-model empirical analysis, this work reveals—for the first time—a pronounced diminishing marginal returns phenomenon in iterative repair: performance gains are concentrated within the first 3–4 iterations, with negligible improvements thereafter. The findings underscore that the design of the repair workflow and feedback mechanisms exerts a far greater influence on repair efficacy than the choice of LLM itself. The authors advocate for treating repair budget as a critical experimental variable to ensure reliable, computationally efficient, and reproducible evaluation outcomes.
This work addresses the high cost and sensitivity to phrasing inherent in manually crafted prompts, as well as the limited ability of existing automated optimization methods to systematically identify and correct failure patterns. To overcome these challenges, the authors propose Reflective Prompt Tuning (RPT), a novel framework that introduces a reflection mechanism leveraging large language model function calling. RPT employs a diagnostic function to analyze failure modes on an optimization set, generates structured reports, and iteratively refines prompts by integrating historical memory with confidence calibration. Experimental results demonstrate that RPT achieves performance gains of up to 12.9 points across three reasoning tasks, with particularly pronounced improvements in multi-hop and mathematical reasoning, while also enhancing the calibration of model output confidence.
This study addresses the common omission in current large language model evaluations of code generation—the iterative refinement process inherent in real-world programming and the models’ capacity for self-correction using feedback. The authors propose a novel framework that leverages execution-based feedback, such as compilation errors and test failures, to systematically investigate how reasoning and non-reasoning models utilize such signals across multiple programming languages. Through multidimensional categorization of code failures and extensive cross-model, cross-language experiments, they demonstrate that reasoning models consistently improve over iterations and significantly outperform non-reasoning counterparts. While syntactic and runtime errors prove relatively amenable to correction, logical and algorithmic errors remain challenging, thereby delineating the current limits of feedback-driven repair mechanisms.
This study presents the first systematic investigation into the calibration of safety confidence in large language models for code generation—specifically, whether a model’s confidence in the safety of its generated code aligns with actual risk. Evaluations are conducted across multiple temperature settings using GPT-4o-mini, Gemini-2.0-Flash, and Qwen3-Coder-Next in both isolated tasks and repository-scale, multi-language scenarios. The work introduces techniques such as calibration-guided repair and architectural gating to improve alignment. Findings reveal that models consistently exhibit overconfidence; safety calibration significantly outperforms functional correctness calibration; calibration-guided repair yields limited gains and often induces functional regressions; and while architectural gating shows promise in controlled settings, its effectiveness degrades in real-world repositories.