llm-driven code generation

Designs and builds systems that use large language models to generate, synthesize, and iteratively refine source code and programs—including agentic or closed‑loop agents that decompose tasks into prompts, invoke toolchains, and run runtime checks. Analyzes and evaluates the produced code with automated tests, semantic code analysis, and large‑scale code mining to measure correctness, behavior, and style and to drive further automated refinement.

llm-drivencodegeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$216K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

A Survey On Large Language Models For Code Generation

Mar 03, 2025
NH
Nam Huynh
🏛️ University of Oklahoma

This paper systematically reviews bottlenecks, optimization strategies, and evaluation frameworks for large language models (LLMs) in code generation. It addresses key limitations in functional correctness, readability, and security. To overcome these, the paper unifies major fine-tuning paradigms—including supervised fine-tuning, instruction tuning, reinforcement learning from human feedback (RLHF), and code-specific pretraining—and proposes a multi-granularity evaluation framework covering syntax, semantics, security, and engineering best practices. Its core contribution is a novel “Capability–Optimization–Evaluation–Deployment” full-stack analytical model that, for the first time, coherently maps the fine-tuning methodology spectrum to industrial tools (e.g., CodeLlama, GitHub Copilot). The work clarifies current technical boundaries, distills reusable methodological guidelines, and provides both theoretical foundations and practical pathways to enable low-barrier, cross-domain programming democratization. (149 words)

LLMs enable non-technical users to generate code using human language.Review of metrics, benchmarks, and applications of LLMs in code generation.Survey explores fine-tuning techniques to improve LLMs' code generation performance.

Must-Read Papers

Most classic and influential ideas
View more

This work proposes an intelligent programming agent that integrates large language models (LLMs) with the GCC compiler to iteratively refine generated code through compilation feedback. The approach addresses the prevalent issue of syntactic and linking errors in LLM-generated code, which often renders it non-compilable. A systematic evaluation across 699 C programming tasks from RosettaCode demonstrates that compiler feedback substantially improves code compilability: using 16 LLMs ranging from 135M to 70B parameters, compilation success rates increase by 5.3 to 79.4 percentage points, syntactic errors decrease by 75%, and undefined reference errors drop by 87%. Notably, smaller models augmented with compiler feedback can outperform larger, unassisted models, highlighting the significant potential of tool-augmented approaches for sustainable and efficient software development.

Code GenerationCompiler IntegrationLarge Language Models

Existing benchmarks primarily assess language models on localized programming tasks, failing to capture their capability to construct complete software systems from scratch. This work introduces ProgramBench, the first end-to-end evaluation framework grounded in behavioral equivalence, which requires agents to autonomously design and implement full codebases based solely on program specifications and documentation, with correctness verified through behavioral test suites. The benchmark encompasses 200 real-world software tasks—including CLI tools, FFmpeg, and SQLite—supports unconstrained, open-ended code generation, and incorporates agent-driven fuzz testing to automatically produce behavioral test cases. Evaluation across nine prominent language models reveals that none can fully solve any task; the best-performing model passes 95% of tests on only 3% of tasks and tends to generate single-file implementations structurally divergent from human-written code.

benchmarkingcode generationlanguage models

This work addresses the challenge that code generated by large language models (LLMs) often suffers from high complexity, redundancy, and architectural debt, and struggles to autonomously identify and perform human-level refactoring. To this end, we introduce CodeTaste, a benchmark that systematically evaluates LLMs’ ability to detect and reproduce real-world refactorings in multi-file settings. Our approach combines large-scale mining of open-source changes, data-flow analysis, and static pattern detection, employing a two-stage “propose-and-implement” strategy. Refactoring quality is validated through test suites and behavioral equivalence checks. Experiments show that while current LLMs can effectively refactor under explicit instructions, they still exhibit a significant gap in autonomously understanding human refactoring intent. Performance is notably enhanced by the staged strategy and by prioritizing proposals aligned with human practices.

behavior-preserving transformationcode qualitycode refactoring

Traditional code review relies heavily on static rule-based checks, lacking proactive risk prediction and pedagogical support. To address this, we propose an intelligent code review agent powered by large language models (LLMs). Our method fine-tunes LLMs on heterogeneous code-related data—including vast codebases, historical review comments, defect reports, and best-practice documentation—and integrates deep code semantic understanding with developer sentiment analysis. The agent performs code smell detection, anticipatory defect identification, actionable improvement suggestions, and educational feedback. Our key contributions are twofold: (1) the first application of LLMs to *proactive* code risk forecasting—moving beyond retrospective rule matching—and (2) an empirically grounded human-AI collaborative evaluation framework. Experimental results demonstrate a statistically significant reduction in post-release defect density, improved review throughput, and high developer acceptance, validating the agent’s dual efficacy in enhancing software quality assurance and fostering engineering skill development.

Detecting code smells, bugs, and providing improvement suggestionsImproving software quality and efficiency through LLM-based code reviewPredicting future potential risks in code to enhance development lifecycle

Exploring the Integration of Large Language Models in Industrial Test Maintenance Processes

Sep 10, 2024
LL
Ludvig Lemner
🏛️ Chalmers University of Technology | Ericsson AB | Chalmers and the University of Gothenburg

To address high manual effort, delayed response, and incomplete coverage in industrial software test maintenance, this paper proposes and empirically validates two novel multi-agent architectures leveraging large language models (LLMs) to accurately identify test cases requiring maintenance—and to assist in their automated addition, deletion, or modification—following code changes. The study systematically distills critical triggering conditions and practical deployment constraints for LLMs in industrial settings, and conducts empirical evaluation using real-world data from Ericsson AB. Results demonstrate significant reductions in manual test maintenance effort, improved timeliness of issue response, and enhanced test coverage completeness. This work represents the first application of multi-agent collaboration to industrial-scale test maintenance, establishing a reusable methodological framework and empirical foundation for deploying LLMs in high-reliability software engineering contexts.

Automating test maintenance to reduce costs and improve qualityExploring LLMs' capabilities for industrial test maintenance supportProposing multi-agent architecture for predicting test maintenance needs

Latest Papers

What's happening recently
View more

There remains a significant gap between academic research on code large language models (Code LLMs) and their industrial deployment. Method: This work systematically investigates the full lifecycle evolution of Code LLMs, establishing a comprehensive technical stack encompassing code pretraining, supervised fine-tuning, reinforcement learning, advanced prompting (e.g., in-context learning, instruction tuning), and autonomous coding agents—all empirically grounded in the Transformer architecture. Contribution/Results: We uncover stage-specific scaling laws, hyperparameter sensitivities, and architectural trade-offs; present the first unified empirical comparison of general-purpose LLMs versus specialized Code LLMs across code correctness, security, and large-codebase contextual awareness. Our models achieve >95% pass@1 on HumanEval and demonstrate practical feasibility on real-world software engineering tasks, delivering a reproducible methodology and implementation paradigm for transitioning code intelligence from research labs to production environments.

Analyzing capabilities of general and specialized code language modelsBridging the research-practice gap in real-world code deploymentSystematically examining the complete lifecycle of code LLMs

Agent-based code generation for the Gammapy framework

Sep 30, 2025
DK
Dmitriy Kostunin
🏛️ Deutsches Elektronen-Synchrotron DESY | JetBrains Limited | JetBrains GmbH

Scientific libraries such as Gammapy suffer from sparse documentation, unstable APIs, and limited community support, severely hindering large language models (LLMs) from generating reliable code. Method: We propose an autonomous code generation agent tailored for scientific frameworks, integrating an LLM, an executable code sandbox, and a lightweight web interface to establish a “generate–execute–verify” closed-loop workflow—reducing reliance on comprehensive documentation and API stability. Contribution/Results: Our approach introduces a dynamic feedback-driven iterative debugging mechanism and a verification protocol specifically designed for scientific computing. We implement a prototype system and a benchmarking suite, empirically validating its effectiveness and scalability on real-world gamma-ray data analysis tasks. Results demonstrate substantial improvements in the robustness and practical utility of LLM-generated code within resource-constrained, domain-specific scientific contexts.

Addressing unstable APIs and limited training data in GammapyCreating controlled environment for writing, executing, and validating codeDeveloping agent-based code generation for specialized scientific libraries

RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring

Nov 05, 2025
KO
Khouloud Oueslati
🏛️ Polytechnique Montreal

Traditional LLMs rely on static instructions for software refactoring, limiting their ability to dynamically adapt to context and make autonomous decisions. This paper introduces RefAgent—the first end-to-end, multi-agent LLM framework for software refactoring—featuring a four-phase collaborative mechanism: planning, execution, testing, and introspective optimization, enabling context-aware refactoring decisions and iterative quality improvement. Its key contributions are: (1) a specialized multi-agent architecture with role-based division of labor; (2) integrated self-reflection and tool-calling capabilities; and (3) a closed-loop, verification-driven paradigm for quality enhancement. Evaluated across eight Java projects, RefAgent achieves a 90% unit test pass rate, reduces code smell median by 52.5%, improves architectural quality attributes by 8.6%, and attains a 79.15 F1-score for refactoring identification—significantly outperforming both single-agent baselines and conventional search-based approaches.

Developing automated software refactoring using multi-agent LLM frameworksEnhancing refactoring accuracy through specialized planning and testing agentsImproving code quality by reducing smells and enhancing key attributes

Hot Scholars

QZ

Qingfu Zhang

Chair Professor, FIEEE, City University of Hong Kong
evolutionary computationmultiobjective optimizationcomputational intelligence
KL

Kevin Leach

Vanderbilt University
Artificial IntelligenceSoftware EngineeringSecurity
PL

Pietro Liguori

Assistant Professor, University of Naples Federico II
AI TrustworthinessLarge Language ModelsCode GenerationFault-Injection
BX

Baowen Xu

Nanjing University
SoftwareProgramming Languages