code generation

Designs, builds, and evaluates systems that automatically produce source code or code fragments from higher-level inputs such as specifications, models, templates, or natural language; this includes program synthesizers, transpilers, compiler backends, and neural or rule-based code generators. Work covers parsing or interpreting the input, mapping to programming constructs, ensuring correctness, readability and performance, and integrating generated code into larger software systems.

codegeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$218K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

SpecGen: Automated Generation of Formal Program Specifications via Large Language Models

Jan 16, 2024
LM
Lezhi Ma
🏛️ Nanjing University | Nanyang Technological University | Singapore Management University

Formal program specifications are notoriously difficult, error-prone, and inefficient to write manually. To address this, we propose a two-stage LLM-driven approach: dialogue-guided specification synthesis followed by mutation-based verification. First, multi-turn dialogues model complex semantic requirements; second, four mutation operators—insertion, replacement, deletion, and reordering—enable verifiability-driven selection, eliminating reliance on rigid templates or syntactic grammars. Our method integrates code understanding, prompt engineering, and heuristic verifiability assessment. Evaluated on SV-COMP and a custom Java benchmark comprising 385 programs, it generates 279 verifiable specifications. These achieve significantly higher completeness and accuracy than pure-LLM baselines and classical tools (e.g., Houdini, Daikon). To our knowledge, this is the first approach to achieve both high coverage and formal verifiability in fully automated specification generation.

Automated generation of formal program specificationsLeveraging LLMs for code comprehensionOvercoming limitations of predefined templates

This work presents the first systematic investigation into the capability of large language models (LLMs) to generate program specifications involving higher-order logical constructs, which are essential for expressing complex verification properties yet remain beyond the reach of existing LLMs that predominantly handle basic syntactic forms. The authors design four syntactic configurations spanning different levels of abstraction and establish a comprehensive evaluation framework to assess a range of representative LLMs on standard verification benchmarks. Experimental results demonstrate that LLMs can effectively produce valid higher-order logical expressions; moreover, integrating logical constructs with base syntax significantly enhances verification efficacy and robustness without substantially increasing verification overhead. The study also reveals distinct advantages of two refinement paradigms in specification generation.

formal specificationlarge language modelslogical constructs

This study evaluates the applicability of large language models (LLMs) to generate **functional and maintainable code** in highly specialized, closed industrial software environments—exemplified by ASML’s ecosystem—facing two key challenges: stringent domain-specific constraints and cross-module code dependencies. We propose a **customized evaluation framework**, introducing the novel metric **build@k**, which quantifies compilation and integration success rates of generated code within real industrial repositories, and establish the first code-generation benchmark tailored to ASML’s proprietary ecosystem. Leveraging few-shot prompting and chain-of-thought strategies, we conduct systematic comparisons across general-purpose and code-specialized LLMs at multiple scales, using both matching-based and execution-based evaluation. Results show that prompt engineering substantially improves build success (few-shot + CoT yields optimal performance), and model specialization benefits exhibit family-level dependency. Our core contributions are an industrial-grade evaluation paradigm for deployable code generation and the first empirically grounded benchmark for this domain.

Assessing code maintainability and integration within specialized software environmentsEvaluating LLM-generated code functionality in industrial proprietary settingsInvestigating impact of prompting techniques and model size on output quality

Large language models (LLMs) exhibit pervasive output formatting bias in code translation tasks—generated outputs frequently contain extraneous natural-language explanations or formatting delimiters, causing standard evaluation metrics (e.g., computation accuracy, CA) to systematically underestimate true performance. Method: We systematically evaluate 11 instruction-tuned LLMs across five programming languages and find that 26.4%–73.7% of translations require post-hoc processing to extract clean code. To address this, we propose a robust code extraction method integrating regex-based parsing with prompt engineering. Contribution/Results: Our approach achieves a 92.73% average Code Extraction Success Rate (CSR) on a multilingual alignment benchmark, substantially improving evaluation fidelity. This work is the first to quantify the impact of formatting bias and establishes a new, generalizable, and robust code extraction paradigm—providing a reproducible, standardized evaluation benchmark for LLM-based code translation.

Evaluating LLM code translation suffers from output format biasesNon-code elements in outputs interfere with performance assessment metricsProposing methods to extract source code for reliable model evaluation

Towards Automated Verification of LLM-Synthesized C Programs

Oct 18, 2024
PM
Prasita Mukherjee
🏛️ Purdue University

Automatically verifying C programs generated by large language models (LLMs) remains challenging due to their syntactic and semantic irregularities, which hinder formal verification. Method: This paper proposes SynVer—a novel framework that tightly integrates LLM-based program synthesis with formal verification. SynVer introduces verifiability-aware biasing mechanisms operating at both syntactic and semantic levels to guide LLMs toward generating verification-friendly code. It further incorporates separation logic (SL) specifications and the Verified Software Toolchain (VST) to enable end-to-end, fully automated verification—from specification to C implementation to machine-checked safety proofs. Results: Evaluated on diverse benchmarks covering basic coding tasks, SL assertions, and API specifications, SynVer significantly improves the automatic verification success rate of LLM-generated C programs. Empirical results demonstrate its scalability, robustness, and effectiveness in bridging the gap between neural code generation and rigorous formal assurance.

Automating verification of LLM-generated C programsDeveloping scalable specification-verification tool for synthesisImposing syntactic and semantic biases for verification

Latest Papers

What's happening recently
View more

Automated generation of verifiable formal specifications is often hindered by syntactic errors, logical inaccuracies, inadequate handling of control-flow structures, and the absence of dynamic error-correction mechanisms. This work proposes AutoReSpec, a novel framework featuring a two-stage collaborative generation mechanism that synergistically combines open- and closed-source large language models. By dynamically selecting prompting strategies based on program structure and invoking a collaborative model upon primary model failure, AutoReSpec leverages feedback from a formal verifier to iteratively refine specifications. Through structure-aware scheduling and a verification-in-the-loop architecture, the approach significantly enhances robustness and efficiency, achieving a 58.2% success rate and 69.2% completeness across 72 Java benchmarks, while reducing average evaluation time by 26.89% compared to existing methods.

formal specification generationLarge Language Modelslogical inaccuracies

This work proposes a human-AI collaborative paradigm for formal software specification that mitigates the traditional barriers to industrial adoption—namely, the notational complexity and high expertise threshold—while preserving the benefits of early error detection and explicit invariants. The approach employs an intermediate language blending natural language with lightweight LaTeX mathematical notation, enabling AI-assisted review, refinement, and code generation. Crucially, it distinguishes between components requiring rigorous formalization and those amenable to flexible treatment. By deeply integrating AI into the specification authoring and verification workflow, this method achieves “correct-by-construction” development in a case study on organizational knowledge growth simulation, significantly reducing costs while ensuring early validation and design correctness.

AI-assisted developmentformal specificationindustrial adoption

This work addresses the limited adoption of formal verification, which often requires expert-written annotations such as preconditions, postconditions, and loop invariants. To overcome this barrier, the authors propose a novel approach that leverages large language models (LLMs) in conjunction with assertions from test cases as static oracles to automatically generate Dafny verification annotations from code annotated with natural language comments. The method features an iterative refinement process guided by verifier feedback over multiple rounds and uniquely integrates multi-model LLM collaboration with a closed-loop verifier feedback mechanism. A VS Code plugin was developed to support practical deployment. Evaluated on 110 Dafny programs, the approach achieves a 98.2% annotation correctness rate within at most eight repair iterations. Empirical results highlight that proof-assistant-style annotation remains a key challenge for LLMs, while user feedback on the plugin was notably positive.

Dafnyformal specificationLLMs

This work addresses a critical limitation in current AI-based code review systems: in the absence of executable specifications, they often fall into structural loops and exhibit correlated errors due to shared training distributions between generation and review models, making it difficult to verify whether code aligns with true intent. The paper proposes a three-tiered architecture—“specification-first, deterministic verification, AI review of residual issues”—and provides the first systematic demonstration that executable specifications can transform code review from a complex domain into a complex yet solvable one. It further clarifies that AI should focus specifically on structural and architectural flaws beyond specification coverage. Through deterministic verification, cross-model review, and targeted defect injection experiments, the study confirms that both intra- and inter-family large models exhibit error correlation without specification guidance, whereas a specification-driven architecture effectively isolates the AI review boundary and significantly enhances reliability.

AI-assisted code reviewcode qualitycorrelated failures

Hot Scholars

YL

Yuyu Luo

Assistant Professor, HKUST(GZ) / HKUST
Data AgentsLLM AgentsDatabaseText-to-SQL
XS

Xiaofei Sun

Stony Brook University, Zhejiang University
Social and Information NetworkNatural Language ProcessingMachine Learning
JT

Julian Togelius

Associate Professor of Computer Science and Engineering, New York University; co-founder, modl.ai
Artificial IntelligenceGamesEvolutionary ComputationGame AI
TH

Torsten Hoefler

Professor of Computer Science at ETH Zurich
High Performance ComputingDeep LearningNetworkingMessage Passing Interface