Score
Designs and builds toolchains that convert script-like inputs (pseudocode, DSLs, or non-executable script files) into runnable code. This includes parsing, code generation or transpilation, dependency resolution, compilation/packaging, and producing artifacts ready for execution on the target runtime.
This paper systematically reviews bottlenecks, optimization strategies, and evaluation frameworks for large language models (LLMs) in code generation. It addresses key limitations in functional correctness, readability, and security. To overcome these, the paper unifies major fine-tuning paradigms—including supervised fine-tuning, instruction tuning, reinforcement learning from human feedback (RLHF), and code-specific pretraining—and proposes a multi-granularity evaluation framework covering syntax, semantics, security, and engineering best practices. Its core contribution is a novel “Capability–Optimization–Evaluation–Deployment” full-stack analytical model that, for the first time, coherently maps the fine-tuning methodology spectrum to industrial tools (e.g., CodeLlama, GitHub Copilot). The work clarifies current technical boundaries, distills reusable methodological guidelines, and provides both theoretical foundations and practical pathways to enable low-barrier, cross-domain programming democratization. (149 words)
This work addresses the challenge of automating library API migration in the absence of real-world migration examples. To overcome this limitation, the authors propose a novel unsupervised approach that leverages large language models (LLMs) to generate initial migration examples without requiring labeled data. These examples are then generalized by an intelligent agent into structured, testable code transformation rules, which are integrated into the PolyglotPiranha framework for execution. This study represents the first integration of LLMs’ zero-shot generation capabilities with programmatic code transformation tools. The method successfully synthesizes reusable and generalizable migration scripts across multiple Python library migration tasks, significantly enhancing the feasibility and practicality of API migration in fully unsupervised settings.
This survey addresses core challenges in code generation for low-resource programming languages (LRPLs) and domain-specific languages (DSLs): severe data scarcity, pronounced syntactic specificity, and inadequate coverage by general-purpose pretraining corpora. We systematically analyze 111 studies published between 2020 and 2024. Methodologically, we propose the first dedicated survey framework for LRPLs/DSLs, categorizing evaluation techniques into four types, performance-enhancement methods into six classes, and identifying emerging adaptation architectures; we further pinpoint the critical absence of standardized benchmarks. Through bibliometric analysis, cross-lingual capability assessment, dataset strategy dissection, and comparative evaluation using multidimensional quality metrics—including CodeBLEU and functional correctness—we empirically delineate the capabilities and limitations of mainstream models (e.g., Codex, CodeLlama). Our contributions include a reusable methodological taxonomy and a practical guideline, establishing foundational support for standardization and future research in LRPL/DSL code generation.
This work addresses the limited cross-file contextual awareness of code large language models (CodeLLMs) in repository-level code generation. Methodologically, we introduce RepoExec—the first executable and functionally correct repository-level benchmark—and propose Dependency Invocation Rate (DIR), a novel metric quantifying the accuracy of cross-file dependency invocation. We further design an instruction-tuning dataset integrating test-driven validation and context-aware dependency modeling. Our contributions include the first comprehensive evaluation framework encompassing context-awareness, execution-driven assessment, and cross-file dependency modeling. Experimental results demonstrate that instruction tuning significantly improves contextual utilization and debugging capability, whereas pre-trained models exhibit stronger functional correctness. RepoExec has since become the de facto standard benchmark for repository-level code generation research.
Existing LLM code-generation evaluations rely heavily on controlled benchmarks (e.g., HumanEval), which poorly reflect real-world development practices. Method: This paper presents the first large-scale empirical study of code generated by ChatGPT and GitHub Copilot on GitHub, integrating repository crawling, language identification, commit-history tracing, complexity measurement, and pattern mining to characterize distribution, evolution, and maintenance properties. Contributions/Results: (1) Quantifies low real-world adoption: LLM-generated code constitutes <1.2% of total codebase volume on average; only 3–8% of such code undergoes modification for bug fixes; and generation is heavily skewed toward Python, Java, and TypeScript. (2) Reveals that associated projects tend to be small-scale, actively evolving, yet critically deficient in documentation. (3) Bridges the gap between controlled benchmarking and engineering practice, establishing an empirical foundation for assessing LLM code trustworthiness and guiding tool optimization.
To address insufficient cross-file context utilization in repository-level code generation—particularly the challenge of balancing general knowledge with fine-grained type dependencies in statically typed languages—this paper proposes a type-dependency-driven context enhancement method. Our approach integrates static-analysis-derived type dependency graphs (for Java and Rust) with multi-file retrieval results to construct semantically richer, structured prompts, thereby overcoming the locality limitations of conventional retrieval methods. The core contribution is the first-ever type-dependency-driven context integration mechanism, enabling principled, cross-file and cross-module modeling of structural knowledge. Evaluated on 199 Java and 90 Rust tasks, our method achieves up to a 17.35% improvement in pass@k over RepoCoder. Crucially, it demonstrates strong generalizability across diverse code-specific and general-purpose large language models.
This study addresses the challenge of generating or modifying industrial-scale, multi-file domain-specific language (DSL) code from natural language instructions using large language models. The work proposes an end-to-end approach that encodes Xtext-based DSL repositories into a path-preserving JSON format, enabling coherent cross-file edits within a single model response. Key innovations include the first demonstration of single-instruction, multi-file DSL generation in industrial settings, a structure-preserving JSON representation, and task-oriented evaluation metrics. Leveraging Qwen2.5-Coder and DeepSeek-Coder (7B) models fine-tuned with QLoRA and augmented by in-context learning, the method achieves perfect structural fidelity (1.00), high exact match accuracy, and strong edit similarity on the test set. Practical utility is further confirmed through developer surveys and successful downstream compilation.
This study investigates the executability of code generated by large language models (LLMs) in clean, minimal environments, revealing a substantial gap between declared dependencies and actual runtime requirements. We propose a novel three-layer dependency framework—comprising declared, available, and runtime dependencies—to systematically quantify dependency inconsistencies in LLM-based programming agents and assess cross-language reproducibility. Using standardized prompt sets across Python, JavaScript, and Java, we conduct automated dependency parsing and environment validation on Claude Code, Codex, and Gemini. Results show that only 68.3% of generated projects execute out-of-the-box; execution success rates are 89.2% for Python and 44.0% for Java, highlighting language-specific disparities. On average, dependency graphs inflate 13.5× relative to declared dependencies, exposing pervasive implicit dependency issues. This work provides critical empirical evidence and a methodological foundation for improving the reliability and engineering deployability of LLM-generated code.
This work addresses the inefficiency of modern JavaScript compilers, which often waste substantial computational resources by indiscriminately applying all downlevel transformations regardless of the actual language features used. To mitigate this, the authors propose a conditional transpilation mechanism that precisely detects and dynamically tracks the set of language features employed at the script level, triggering transformations only for those features that require them. Implemented within the Google Closure Compiler, the approach integrates feature set construction, strategic pass ordering, and post-transpilation feature validation to significantly reduce unnecessary abstract syntax tree (AST) traversals. Empirical evaluation on large-scale production codebases demonstrates that the proposed method effectively decreases compilation time while reducing both memory consumption and computational overhead.
Manually configuring linters requires expert knowledge and struggles to adapt across multiple programming languages, coding standards, and tooling ecosystems, leading to high maintenance overhead. This work proposes LintCFG, the first approach to apply compiler design principles to automated linter configuration generation. It introduces a tool-agnostic domain-specific language (DSL) to structurally encode coding rules and leverages large language models to automatically compile natural language specifications into concrete linter configurations, enabling end-to-end automation across languages, standards, and tools. Evaluated on Java Checkstyle tasks, the DSL achieves over 90% precision and recall in rule representation, with fine-grained configuration generation exceeding 70% accuracy—more than doubling the performance of baseline methods. User studies confirm significant gains in developer productivity, and the approach successfully generalizes to JavaScript ESLint scenarios.
Automatically generating verifiable Python formal specifications remains challenging, and developers often abandon automated verification tools due to the tediousness of manually writing contracts. This work proposes a closed-loop approach that integrates large language models with symbolic execution (CrossHair) to automatically generate and iteratively refine icontract-style contract annotations without modifying the original code. The method leverages feedback from symbolic execution to drive specification refinement and simultaneously produces coverage-guided pytest stubs and debugging artifacts. Experimental results demonstrate that the approach successfully generates CrossHair-compatible specifications for most programs, significantly enhancing the practical feasibility of automated verification, while also revealing real-world limitations arising from the boundaries of symbolic exploration and behavioral discrepancies in large language models.