From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence

πŸ“… 2025-11-23
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
There remains a significant gap between academic research on code large language models (Code LLMs) and their industrial deployment. Method: This work systematically investigates the full lifecycle evolution of Code LLMs, establishing a comprehensive technical stack encompassing code pretraining, supervised fine-tuning, reinforcement learning, advanced prompting (e.g., in-context learning, instruction tuning), and autonomous coding agentsβ€”all empirically grounded in the Transformer architecture. Contribution/Results: We uncover stage-specific scaling laws, hyperparameter sensitivities, and architectural trade-offs; present the first unified empirical comparison of general-purpose LLMs versus specialized Code LLMs across code correctness, security, and large-codebase contextual awareness. Our models achieve >95% pass@1 on HumanEval and demonstrate practical feasibility on real-world software engineering tasks, delivering a reproducible methodology and implementation paradigm for transitioning code intelligence from research labs to production environments.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
πŸ“ Abstract
Large language models (LLMs) have fundamentally transformed automated software development by enabling direct translation of natural language descriptions into functional code, driving commercial adoption through tools like Github Copilot (Microsoft), Cursor (Anysphere), Trae (ByteDance), and Claude Code (Anthropic). While the field has evolved dramatically from rule-based systems to Transformer-based architectures, achieving performance improvements from single-digit to over 95% success rates on benchmarks like HumanEval. In this work, we provide a comprehensive synthesis and practical guide (a series of analytic and probing experiments) about code LLMs, systematically examining the complete model life cycle from data curation to post-training through advanced prompting paradigms, code pre-training, supervised fine-tuning, reinforcement learning, and autonomous coding agents. We analyze the code capability of the general LLMs (GPT-4, Claude, LLaMA) and code-specialized LLMs (StarCoder, Code LLaMA, DeepSeek-Coder, and QwenCoder), critically examining the techniques, design decisions, and trade-offs. Further, we articulate the research-practice gap between academic research (e.g., benchmarks and tasks) and real-world deployment (e.g., software-related code tasks), including code correctness, security, contextual awareness of large codebases, and integration with development workflows, and map promising research directions to practical needs. Last, we conduct a series of experiments to provide a comprehensive analysis of code pre-training, supervised fine-tuning, and reinforcement learning, covering scaling law, framework selection, hyperparameter sensitivity, model architectures, and dataset comparisons.
Problem

Research questions and friction points this paper is trying to address.

Systematically examining the complete lifecycle of code LLMs
Analyzing capabilities of general and specialized code language models
Bridging the research-practice gap in real-world code deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Systematic model lifecycle from data curation to deployment
Advanced prompting paradigms and autonomous coding agents
Comprehensive analysis of pre-training and fine-tuning techniques
J
Jian Yang
Beihang University, M-A-P
W
Wei Zhang
Beihang University, M-A-P
S
Shark Liu
Beihang University, M-A-P
J
Jiajun Wu
Beihang University, M-A-P
S
Shawn Guo
Beihang University, M-A-P
Yizhi Li
Yizhi Li
University of Manchester, M-A-P
LLMReasoningPost-trainingComputational Music