Score
Designs, builds, and evaluates systems that automatically detect and diagnose runtime and dependency failures, synthesize corrective code patches, and apply or revert those patches to restore reproducible execution. These systems implement execution-guided workflows — e.g., generate-run-revise loops, staged or guarded re-execution, execution-feedback and test-based validation, prioritization of fixes, memory or state interventions, and optional LLM guidance — to validate, refine, and iterate candidate repairs until robust behavior is achieved.
This study systematically reviews 63 LLM-based automated program repair (APR) systems published between January 2022 and June 2025, addressing three core challenges: semantic correctness verification beyond test suites, large-scale repository-level defect repair, and optimization of LLM inference cost. We propose the first comprehensive taxonomy, categorizing APR designs into four paradigms: fine-tuning, prompt engineering, pipeline-based workflows, and agent frameworks. Quantitative analysis demonstrates how retrieval augmentation and static/dynamic code analysis enhance context quality, while revealing fundamental trade-offs among cost, controllability, and scalability across paradigms. Key insights identify lightweight feedback mechanisms, repository-aware retrieval, and cost-aware planning as critical levers for advancement. The work establishes a theoretical framework and practical roadmap to enhance the reliability, scalability, and real-world applicability of LLM-APR systems.
This study addresses the lack of systematic understanding regarding the impact of repair loop iteration counts in large language model (LLM)-based software engineering tasks, where prior work often relies on arbitrarily defined repair budgets. Through a cross-task (code generation, test generation, code translation) and cross-model empirical analysis, this work reveals—for the first time—a pronounced diminishing marginal returns phenomenon in iterative repair: performance gains are concentrated within the first 3–4 iterations, with negligible improvements thereafter. The findings underscore that the design of the repair workflow and feedback mechanisms exerts a far greater influence on repair efficacy than the choice of LLM itself. The authors advocate for treating repair budget as a critical experimental variable to ensure reliable, computationally efficient, and reproducible evaluation outcomes.
Existing automated program repair techniques struggle to address complex logical errors and silent failures due to their inability to accurately model runtime dynamic behaviors and data dependencies. This work proposes TraceRepair, a novel framework that, for the first time, incorporates runtime execution traces as shared constraints within a multi-agent collaboration mechanism. In this approach, a probe agent captures snapshots of critical variables, while multiple specialized agents—powered by large language models—perform cross-validation and iterative refinement to enable precise, dynamic-reasoning-driven repairs. Evaluated on Defects4J, TraceRepair successfully fixes 392 bugs, substantially outperforming current LLM-based methods, and demonstrates strong generalization capabilities on a newly curated dataset of recent vulnerabilities.
Large language models can generate runnable software artifacts, but their security remains difficult to evaluate end to end. This study examines that problem through a Detect--Repair--Verify (DRV) workflow, in which vulnerabilities are detected, repaired, and then rechecked with security and functional tests. It addresses four gaps in current evidence: the lack of test-grounded benchmarks for LLM-generated artifacts, limited evidence on pipeline-level effectiveness, unclear reliability of detection reports as repair guidance, and uncertain repair trustworthiness under verification. To support this study, EduCollab is constructed as a multi-language, multi-granularity benchmark of runnable LLM-generated web applications in PHP, JavaScript, and Python. Each artifact is paired with executable functional and exploit test suites, and the benchmark spans project-, requirement-, and file-level settings. On this benchmark, the study compares unrepaired baselines, single-pass detect--repair, and bounded iterative DRV under comparable budget constraints. Outcomes are measured by secure-and-correct yield, and intermediate artifacts and iteration traces are analyzed to assess report actionability and repair failure modes. The results show that bounded iterative DRV can improve secure-and-correct yield over single-pass repair, but the gains are uneven at the project level and become clearer at narrower repair scopes. Detection reports are often useful for downstream repair, but their reliability is inconsistent. Repair trustworthiness also depends strongly on repair scope. These findings highlight the need for test-grounded, end-to-end evaluation of LLM-based vulnerability management workflows.
This work addresses the insufficient correctness guarantees in program synthesis by proposing a collaborative synthesis framework integrating dynamic multi-agent workflows with an LLM-based quality checker. Methodologically, it establishes a closed-loop collaboration among code generation, test execution, and self-debugging agents, and introduces the first LLM quality checker that explicitly models program execution traces to assess test compliance in real time—enabling dynamic submission, issue clarification, and step-level backtracking—augmented by diverse prompting and quality-feedback-driven adaptive decision-making. The key contribution is the first integration of dynamic execution-aware quality verification into the synthesis pipeline, enabling fine-grained procedural control. Empirically, the approach achieves state-of-the-art performance on MBPP, HumanEval, and EvalPlus, significantly outperforming static workflows and zero-shot one-shot synthesis baselines.
This study addresses a critical yet previously underexplored issue in large language model (LLM)-driven software development: the contamination of automatically generated tests by erroneous code. The authors systematically uncover and empirically validate this error propagation phenomenon, demonstrating that when tests are generated based on incorrect code within multi-step agent workflows—across diverse programming tasks and various prompting strategies, including chain-of-thought—the resulting tests exhibit significantly lower defect detection rates (14%) compared to independently generated tests (25%). These findings challenge the prevailing assumption that LLM-generated tests can serve as reliable, independent oracles, thereby highlighting the substantial risk of test bias in LLM-augmented development pipelines.
This work addresses the challenge that test inputs generated by large language models (LLMs) often fail to reliably cover target code paths due to limited contextual awareness and uncontrolled reasoning. To overcome this limitation, the authors propose ReDig, a novel framework that introduces runtime value feedback into LLM-driven test generation for the first time, establishing a closed-loop optimization architecture. By monitoring execution traces of previously failed tests, ReDig diagnoses the causes of path non-reachability and iteratively refines subsequent test scripts. Experimental results on Poppler and Libsndfile demonstrate that ReDig effectively leverages runtime feedback to enhance coverage of targeted code lines, significantly improving the reachability and reliability of LLM-generated tests.
This work addresses a critical yet overlooked reliability issue in code generated by large language models (LLMs): despite passing compilation and unit tests, such code often fails in deployment due to structural inconsistencies—such as missing configurations, invalid imports, or omitted security controls—that evade detection by conventional CI/SAST tools. The paper introduces the “patchwork problem” to characterize these cross-module global defects, proposes an eight-category taxonomy specific to LLM-generated code, and formalizes structural consistency via invariants derived from a multidimensional code graph encompassing imports, calls, dependencies, configurations, and routing. Building on this foundation, the authors design a hybrid verification framework that integrates traditional static analysis with custom graph-based invariant checkers to precisely identify structural flaws invisible to existing tools. Empirical evaluation reveals that such defects are pervasive across major LLMs under diverse prompting strategies and exhibit distinct model-specific patterns.
This work addresses the unreliability of large language model (LLM)-driven agent workflows, which stems from output nondeterminism, complex node dependencies, and tool heterogeneity, and proposes FlowFixer—a novel framework that introduces symbolic reasoning into automated workflow repair. FlowFixer models execution traces symbolically to generate behavioral specifications, enabling precise fault localization and root cause identification, and dynamically synthesizes targeted repair patches. To reduce verification overhead, it incorporates a multidimensional pre-evaluation mechanism. Experimental evaluation on Dify, Coze, and n8n platforms demonstrates that FlowFixer achieves a repair success rate of 71.3%, outperforming existing methods by 11.9%–27.6%, and improves root cause analysis accuracy by 15.3%–38.8%.