Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey

📅 2026-01-15
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) still face significant challenges in automatically repairing software defects in complex, real-world scenarios. This work presents a systematic survey of LLM-driven automated program repair techniques, introducing a novel modular agent framework to categorize existing approaches and establishing a dynamic open-source repository to foster community advancement. The study integrates key methodologies—including automated data synthesis, training-free modular architectures, supervised fine-tuning, and reinforcement learning—and provides an in-depth evaluation of their performance on benchmarks such as SWE-bench. It reveals the critical influence of data quality and agent behavior on repair efficacy, thereby delineating promising directions for future research in the field.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Planning, Routing, and Scheduling: Replanning and Plan RepairNatural Language Processing: Code Generation / Program Synthesis from Natural Language

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Issue resolution, a complex Software Engineering (SWE) task integral to real-world development, has emerged as a compelling challenge for artificial intelligence. The establishment of benchmarks like SWE-bench revealed this task as profoundly difficult for large language models, thereby significantly accelerating the evolution of autonomous coding agents. This paper presents a systematic survey of this emerging domain. We begin by examining data construction pipelines, covering automated collection and synthesis approaches. We then provide a comprehensive analysis of methodologies, spanning training-free frameworks with their modular components to training-based techniques, including supervised fine-tuning and reinforcement learning. Subsequently, we discuss critical analyses of data quality and agent behavior, alongside practical applications. Finally, we identify key challenges and outline promising directions for future research. An open-source repository is maintained at https://github.com/DeepSoftwareAnalytics/Awesome-Issue-Resolution to serve as a dynamic resource in this field.
Problem

Research questions and friction points this paper is trying to address.

issue resolution
large language models
software engineering
autonomous coding agents
SWE-bench
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-based issue resolution
autonomous coding agents
software engineering
data synthesis
agent frameworks
🔎 Similar Papers
No similar papers found.