🤖 AI Summary
This survey addresses the lack of a unified taxonomy and technical landscape in large language model (LLM)-based code generation agent research. To resolve this, we propose a three-dimensional analytical framework centered on autonomy, task scalability, and engineering practicality—enabling the first structured classification of single- and multi-agent architectures and mapping their capabilities across software development lifecycle (SDLC) phases. We systematically review prevailing evaluation benchmarks (e.g., CodeSearchNet, SWE-bench), toolchains (e.g., CodeAgent, DevInfer), and persistent implementation challenges. The study yields the first comprehensive landscape integrating technical foundations, application paradigms, and evaluation methodologies. It identifies key long-term research directions, including verifiable autonomy, robust planning under uncertainty, and human–AI collaborative workflow modeling—thereby establishing a foundational reference for advancing LLM-powered code agents.
📝 Abstract
Code generation agents powered by large language models (LLMs) are revolutionizing the software development paradigm. Distinct from previous code generation techniques, code generation agents are characterized by three core features. 1) Autonomy: the ability to independently manage the entire workflow, from task decomposition to coding and debugging. 2) Expanded task scope: capabilities that extend beyond generating code snippets to encompass the full software development lifecycle (SDLC). 3) Enhancement of engineering practicality: a shift in research emphasis from algorithmic innovation toward practical engineering challenges, such as system reliability, process management, and tool integration. This domain has recently witnessed rapid development and an explosion in research, demonstrating significant application potential. This paper presents a systematic survey of the field of LLM-based code generation agents. We trace the technology's developmental trajectory from its inception and systematically categorize its core techniques, including both single-agent and multi-agent architectures. Furthermore, this survey details the applications of LLM-based agents across the full SDLC, summarizes mainstream evaluation benchmarks and metrics, and catalogs representative tools. Finally, by analyzing the primary challenges, we identify and propose several foundational, long-term research directions for the future work of the field.