Score
Designs and builds systems that use large language models to generate, synthesize, and iteratively refine source code and programs—including agentic or closed‑loop agents that decompose tasks into prompts, invoke toolchains, and run runtime checks. Analyzes and evaluates the produced code with automated tests, semantic code analysis, and large‑scale code mining to measure correctness, behavior, and style and to drive further automated refinement.
This survey addresses the lack of a unified taxonomy and technical landscape in large language model (LLM)-based code generation agent research. To resolve this, we propose a three-dimensional analytical framework centered on autonomy, task scalability, and engineering practicality—enabling the first structured classification of single- and multi-agent architectures and mapping their capabilities across software development lifecycle (SDLC) phases. We systematically review prevailing evaluation benchmarks (e.g., CodeSearchNet, SWE-bench), toolchains (e.g., CodeAgent, DevInfer), and persistent implementation challenges. The study yields the first comprehensive landscape integrating technical foundations, application paradigms, and evaluation methodologies. It identifies key long-term research directions, including verifiable autonomy, robust planning under uncertainty, and human–AI collaborative workflow modeling—thereby establishing a foundational reference for advancing LLM-powered code agents.
This paper systematically reviews bottlenecks, optimization strategies, and evaluation frameworks for large language models (LLMs) in code generation. It addresses key limitations in functional correctness, readability, and security. To overcome these, the paper unifies major fine-tuning paradigms—including supervised fine-tuning, instruction tuning, reinforcement learning from human feedback (RLHF), and code-specific pretraining—and proposes a multi-granularity evaluation framework covering syntax, semantics, security, and engineering best practices. Its core contribution is a novel “Capability–Optimization–Evaluation–Deployment” full-stack analytical model that, for the first time, coherently maps the fine-tuning methodology spectrum to industrial tools (e.g., CodeLlama, GitHub Copilot). The work clarifies current technical boundaries, distills reusable methodological guidelines, and provides both theoretical foundations and practical pathways to enable low-barrier, cross-domain programming democratization. (149 words)
This work proposes an intelligent programming agent that integrates large language models (LLMs) with the GCC compiler to iteratively refine generated code through compilation feedback. The approach addresses the prevalent issue of syntactic and linking errors in LLM-generated code, which often renders it non-compilable. A systematic evaluation across 699 C programming tasks from RosettaCode demonstrates that compiler feedback substantially improves code compilability: using 16 LLMs ranging from 135M to 70B parameters, compilation success rates increase by 5.3 to 79.4 percentage points, syntactic errors decrease by 75%, and undefined reference errors drop by 87%. Notably, smaller models augmented with compiler feedback can outperform larger, unassisted models, highlighting the significant potential of tool-augmented approaches for sustainable and efficient software development.
Existing benchmarks primarily assess language models on localized programming tasks, failing to capture their capability to construct complete software systems from scratch. This work introduces ProgramBench, the first end-to-end evaluation framework grounded in behavioral equivalence, which requires agents to autonomously design and implement full codebases based solely on program specifications and documentation, with correctness verified through behavioral test suites. The benchmark encompasses 200 real-world software tasks—including CLI tools, FFmpeg, and SQLite—supports unconstrained, open-ended code generation, and incorporates agent-driven fuzz testing to automatically produce behavioral test cases. Evaluation across nine prominent language models reveals that none can fully solve any task; the best-performing model passes 95% of tests on only 3% of tasks and tends to generate single-file implementations structurally divergent from human-written code.
This work addresses the challenge that code generated by large language models (LLMs) often suffers from high complexity, redundancy, and architectural debt, and struggles to autonomously identify and perform human-level refactoring. To this end, we introduce CodeTaste, a benchmark that systematically evaluates LLMs’ ability to detect and reproduce real-world refactorings in multi-file settings. Our approach combines large-scale mining of open-source changes, data-flow analysis, and static pattern detection, employing a two-stage “propose-and-implement” strategy. Refactoring quality is validated through test suites and behavioral equivalence checks. Experiments show that while current LLMs can effectively refactor under explicit instructions, they still exhibit a significant gap in autonomously understanding human refactoring intent. Performance is notably enhanced by the staged strategy and by prioritizing proposals aligned with human practices.
Traditional code review relies heavily on static rule-based checks, lacking proactive risk prediction and pedagogical support. To address this, we propose an intelligent code review agent powered by large language models (LLMs). Our method fine-tunes LLMs on heterogeneous code-related data—including vast codebases, historical review comments, defect reports, and best-practice documentation—and integrates deep code semantic understanding with developer sentiment analysis. The agent performs code smell detection, anticipatory defect identification, actionable improvement suggestions, and educational feedback. Our key contributions are twofold: (1) the first application of LLMs to *proactive* code risk forecasting—moving beyond retrospective rule matching—and (2) an empirically grounded human-AI collaborative evaluation framework. Experimental results demonstrate a statistically significant reduction in post-release defect density, improved review throughput, and high developer acceptance, validating the agent’s dual efficacy in enhancing software quality assurance and fostering engineering skill development.
To address high manual effort, delayed response, and incomplete coverage in industrial software test maintenance, this paper proposes and empirically validates two novel multi-agent architectures leveraging large language models (LLMs) to accurately identify test cases requiring maintenance—and to assist in their automated addition, deletion, or modification—following code changes. The study systematically distills critical triggering conditions and practical deployment constraints for LLMs in industrial settings, and conducts empirical evaluation using real-world data from Ericsson AB. Results demonstrate significant reductions in manual test maintenance effort, improved timeliness of issue response, and enhanced test coverage completeness. This work represents the first application of multi-agent collaboration to industrial-scale test maintenance, establishing a reusable methodological framework and empirical foundation for deploying LLMs in high-reliability software engineering contexts.
There remains a significant gap between academic research on code large language models (Code LLMs) and their industrial deployment. Method: This work systematically investigates the full lifecycle evolution of Code LLMs, establishing a comprehensive technical stack encompassing code pretraining, supervised fine-tuning, reinforcement learning, advanced prompting (e.g., in-context learning, instruction tuning), and autonomous coding agents—all empirically grounded in the Transformer architecture. Contribution/Results: We uncover stage-specific scaling laws, hyperparameter sensitivities, and architectural trade-offs; present the first unified empirical comparison of general-purpose LLMs versus specialized Code LLMs across code correctness, security, and large-codebase contextual awareness. Our models achieve >95% pass@1 on HumanEval and demonstrate practical feasibility on real-world software engineering tasks, delivering a reproducible methodology and implementation paradigm for transitioning code intelligence from research labs to production environments.
Scientific libraries such as Gammapy suffer from sparse documentation, unstable APIs, and limited community support, severely hindering large language models (LLMs) from generating reliable code. Method: We propose an autonomous code generation agent tailored for scientific frameworks, integrating an LLM, an executable code sandbox, and a lightweight web interface to establish a “generate–execute–verify” closed-loop workflow—reducing reliance on comprehensive documentation and API stability. Contribution/Results: Our approach introduces a dynamic feedback-driven iterative debugging mechanism and a verification protocol specifically designed for scientific computing. We implement a prototype system and a benchmarking suite, empirically validating its effectiveness and scalability on real-world gamma-ray data analysis tasks. Results demonstrate substantial improvements in the robustness and practical utility of LLM-generated code within resource-constrained, domain-specific scientific contexts.
Traditional LLMs rely on static instructions for software refactoring, limiting their ability to dynamically adapt to context and make autonomous decisions. This paper introduces RefAgent—the first end-to-end, multi-agent LLM framework for software refactoring—featuring a four-phase collaborative mechanism: planning, execution, testing, and introspective optimization, enabling context-aware refactoring decisions and iterative quality improvement. Its key contributions are: (1) a specialized multi-agent architecture with role-based division of labor; (2) integrated self-reflection and tool-calling capabilities; and (3) a closed-loop, verification-driven paradigm for quality enhancement. Evaluated across eight Java projects, RefAgent achieves a 90% unit test pass rate, reduces code smell median by 52.5%, improves architectural quality attributes by 8.6%, and attains a 79.15 F1-score for refactoring identification—significantly outperforming both single-agent baselines and conventional search-based approaches.