π€ AI Summary
There remains a significant gap between academic research on code large language models (Code LLMs) and their industrial deployment. Method: This work systematically investigates the full lifecycle evolution of Code LLMs, establishing a comprehensive technical stack encompassing code pretraining, supervised fine-tuning, reinforcement learning, advanced prompting (e.g., in-context learning, instruction tuning), and autonomous coding agentsβall empirically grounded in the Transformer architecture. Contribution/Results: We uncover stage-specific scaling laws, hyperparameter sensitivities, and architectural trade-offs; present the first unified empirical comparison of general-purpose LLMs versus specialized Code LLMs across code correctness, security, and large-codebase contextual awareness. Our models achieve >95% pass@1 on HumanEval and demonstrate practical feasibility on real-world software engineering tasks, delivering a reproducible methodology and implementation paradigm for transitioning code intelligence from research labs to production environments.
π Abstract
Large language models (LLMs) have fundamentally transformed automated software development by enabling direct translation of natural language descriptions into functional code, driving commercial adoption through tools like Github Copilot (Microsoft), Cursor (Anysphere), Trae (ByteDance), and Claude Code (Anthropic). While the field has evolved dramatically from rule-based systems to Transformer-based architectures, achieving performance improvements from single-digit to over 95% success rates on benchmarks like HumanEval. In this work, we provide a comprehensive synthesis and practical guide (a series of analytic and probing experiments) about code LLMs, systematically examining the complete model life cycle from data curation to post-training through advanced prompting paradigms, code pre-training, supervised fine-tuning, reinforcement learning, and autonomous coding agents. We analyze the code capability of the general LLMs (GPT-4, Claude, LLaMA) and code-specialized LLMs (StarCoder, Code LLaMA, DeepSeek-Coder, and QwenCoder), critically examining the techniques, design decisions, and trade-offs. Further, we articulate the research-practice gap between academic research (e.g., benchmarks and tasks) and real-world deployment (e.g., software-related code tasks), including code correctness, security, contextual awareness of large codebases, and integration with development workflows, and map promising research directions to practical needs. Last, we conduct a series of experiments to provide a comprehensive analysis of code pre-training, supervised fine-tuning, and reinforcement learning, covering scaling law, framework selection, hyperparameter sensitivity, model architectures, and dataset comparisons.