A Survey on Code Generation with LLM-based Agents

📅 2025-07-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This survey addresses the lack of a unified taxonomy and technical landscape in large language model (LLM)-based code generation agent research. To resolve this, we propose a three-dimensional analytical framework centered on autonomy, task scalability, and engineering practicality—enabling the first structured classification of single- and multi-agent architectures and mapping their capabilities across software development lifecycle (SDLC) phases. We systematically review prevailing evaluation benchmarks (e.g., CodeSearchNet, SWE-bench), toolchains (e.g., CodeAgent, DevInfer), and persistent implementation challenges. The study yields the first comprehensive landscape integrating technical foundations, application paradigms, and evaluation methodologies. It identifies key long-term research directions, including verifiable autonomy, robust planning under uncertainty, and human–AI collaborative workflow modeling—thereby establishing a foundational reference for advancing LLM-powered code agents.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageCognitive Modeling & Cognitive Systems: Agent ArchitecturesMachine Learning: Large Multimodal Models (LMMs)

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
Code generation agents powered by large language models (LLMs) are revolutionizing the software development paradigm. Distinct from previous code generation techniques, code generation agents are characterized by three core features. 1) Autonomy: the ability to independently manage the entire workflow, from task decomposition to coding and debugging. 2) Expanded task scope: capabilities that extend beyond generating code snippets to encompass the full software development lifecycle (SDLC). 3) Enhancement of engineering practicality: a shift in research emphasis from algorithmic innovation toward practical engineering challenges, such as system reliability, process management, and tool integration. This domain has recently witnessed rapid development and an explosion in research, demonstrating significant application potential. This paper presents a systematic survey of the field of LLM-based code generation agents. We trace the technology's developmental trajectory from its inception and systematically categorize its core techniques, including both single-agent and multi-agent architectures. Furthermore, this survey details the applications of LLM-based agents across the full SDLC, summarizes mainstream evaluation benchmarks and metrics, and catalogs representative tools. Finally, by analyzing the primary challenges, we identify and propose several foundational, long-term research directions for the future work of the field.
Problem

Research questions and friction points this paper is trying to address.

Autonomous code generation agents manage full development workflow
LLM-based agents expand beyond snippets to full SDLC
Shift focus from algorithms to practical engineering challenges
Innovation

Methods, ideas, or system contributions that make the work stand out.

Autonomous workflow management in code generation
Full software lifecycle task capabilities
Focus on engineering practicality and reliability
🔎 Similar Papers
No similar papers found.
Yihong Dong
Yihong Dong
Peking University
Code GenerationLarge Language Models
X
Xue Jiang
School of Computer Science, Peking University, Beijing, China
J
Jiaru Qian
School of Computer Science, Peking University, Beijing, China
T
Tian Wang
School of Computer Science, Peking University, Beijing, China
Kechi Zhang
Kechi Zhang
Peking University
AI4SE
Zhi Jin
Zhi Jin
Sun Yat-Sen University, Associate Professor
Ge Li
Ge Li
Full Professor of Computer Science, Peking University
Program AnalysisProgram GenerationDeep Learning