automated code transformation

Design, build, or analyze systems that automatically generate, select, instantiate, sequence, and apply code-level transformations (refactorings or context-sensitive rewrites) across software repositories. This work covers specification and matching of transformation patterns, context-sensitive application and sequencing to preserve semantics, and evaluation of downstream effects on program analyses and runtime performance.

automatedcodetransformation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

On the Impacts of Contexts on Repository-Level Code Generation

Jun 17, 2024
NL
Nam Le Hai
🏛️ FPT Software

This work addresses the limited cross-file contextual awareness of code large language models (CodeLLMs) in repository-level code generation. Methodologically, we introduce RepoExec—the first executable and functionally correct repository-level benchmark—and propose Dependency Invocation Rate (DIR), a novel metric quantifying the accuracy of cross-file dependency invocation. We further design an instruction-tuning dataset integrating test-driven validation and context-aware dependency modeling. Our contributions include the first comprehensive evaluation framework encompassing context-awareness, execution-driven assessment, and cross-file dependency modeling. Experimental results demonstrate that instruction tuning significantly improves contextual utilization and debugging capability, whereas pre-trained models exhibit stronger functional correctness. RepoExec has since become the de facto standard benchmark for repository-level code generation research.

Develop reliable benchmarks for CodeLLMsEnhance CodeLLMs' context utilizationEvaluate repository-level code generation

Existing direct code-to-code transformation approaches often suffer from semantic drift, implicit behavioral changes, and loss of traceability. To address these issues, this work proposes a specification-based Code2Text2Code refactoring framework that first translates source code into a neutral textual specification before generating target code. The approach integrates abstract syntax tree (AST) and dependency graph analysis, semantic-aware code chunking, retrieval-augmented generation, and DSPy-based prompt tuning, further enhanced by iterative validation and graph-based formal verification. This pipeline ensures high-fidelity semantic preservation and controllable evolution during code transformation. Experimental results demonstrate that the proposed method significantly reduces transformation loss and substantially improves semantic consistency, interface stability, and cross-language traceability of the refactored code.

behavioral changesCode2Code transformationdomain logic reconstruction

This work addresses the high cost, inconsistency, and poor reproducibility associated with manual collection of method-level contextual information—such as class metadata, documentation, and call relationships—in large-scale Java projects. To overcome these challenges, the authors propose the first task-agnostic, reusable unified pipeline that automatically parses Maven/Gradle project structures and classpaths, leverages SootUp to construct static call graphs, employs Spoon for source code analysis, and achieves precise alignment between source code and bytecode to generate a versioned, multidimensional context dataset. Evaluated on 20 real-world repositories, the pipeline successfully processes 56,512 methods and 386,048 call edges, with 97.8% of intra-project call edges accurately mapped to source code locations and a human-audited correctness rate of 99.0%.

automated dataset generationcode contextJava projects

To address insufficient cross-file context utilization in repository-level code generation—particularly the challenge of balancing general knowledge with fine-grained type dependencies in statically typed languages—this paper proposes a type-dependency-driven context enhancement method. Our approach integrates static-analysis-derived type dependency graphs (for Java and Rust) with multi-file retrieval results to construct semantically richer, structured prompts, thereby overcoming the locality limitations of conventional retrieval methods. The core contribution is the first-ever type-dependency-driven context integration mechanism, enabling principled, cross-file and cross-module modeling of structural knowledge. Evaluated on 199 Java and 90 Rust tasks, our method achieves up to a 17.35% improvement in pass@k over RepoCoder. Crucially, it demonstrates strong generalizability across diverse code-specific and general-purpose large language models.

Existing retrieval methods fail to capture comprehensive type dependenciesLLMs need better context integration for statically typed languagesRepository-level code generation struggles with multi-file context integration

SpecGen: Automated Generation of Formal Program Specifications via Large Language Models

Jan 16, 2024
LM
Lezhi Ma
🏛️ Nanjing University | Nanyang Technological University | Singapore Management University

Formal program specifications are notoriously difficult, error-prone, and inefficient to write manually. To address this, we propose a two-stage LLM-driven approach: dialogue-guided specification synthesis followed by mutation-based verification. First, multi-turn dialogues model complex semantic requirements; second, four mutation operators—insertion, replacement, deletion, and reordering—enable verifiability-driven selection, eliminating reliance on rigid templates or syntactic grammars. Our method integrates code understanding, prompt engineering, and heuristic verifiability assessment. Evaluated on SV-COMP and a custom Java benchmark comprising 385 programs, it generates 279 verifiable specifications. These achieve significantly higher completeness and accuracy than pure-LLM baselines and classical tools (e.g., Houdini, Daikon). To our knowledge, this is the first approach to achieve both high coverage and formal verifiability in fully automated specification generation.

Automated generation of formal program specificationsLeveraging LLMs for code comprehensionOvercoming limitations of predefined templates

Latest Papers

What's happening recently
View more

Existing refactoring tools struggle to automatically identify opportunities to replace custom logic with equivalent API invocations, as they rely on predefined templates and fail to effectively model semantic equivalence across multiple statements. This work presents the first systematic characterization of the scope, categories, and patterns of API-replacement refactoring and introduces AKIRA, a hybrid recommendation framework that integrates static pattern matching with semantic reasoning, augmented by a refactoring-aware knowledge base. Evaluated through empirical analysis of 166,299 open-source Java commits and manual validation, AKIRA achieves 90% recall and 88% precision on an internal dataset and substantially improves performance on the external RETIWA dataset, increasing recall from 21% to 81% and precision from 40% to 78%, thereby significantly enhancing the accuracy and feasibility of identifying complex API-replacement refactorings.

API replacement refactoringcode refactoringcustom logic

Existing benchmarks struggle to evaluate large language models’ ability to adapt code in the absence of explicit instructions, across multiple change types, and at the fragment-level granularity. This work proposes a mutation-injection framework based on open-source Java code, introducing— for the first time at the fragment level—a taxonomy of adaptation operations inspired by real developer behaviors. By leveraging controlled mutations and reinserting test suites, the framework assesses models’ contextual adaptation capabilities without requiring edit instructions. The approach supports multi-granular context control, enabling quantitative analysis of how different adaptation types affect model performance and revealing fundamental limits in scalability with respect to code complexity and contextual dependency.

code adaptationinstruction-freeJava code snippets

Existing repository-level code generation methods struggle to capture functions that share similar procedural logic but differ in identifiers or domains, and they inadequately model cross-file dependencies and project-specific conventions. This work introduces procedural similarity as an explicit retrieval signal, decomposing the target function into intermediate reasoning steps. At each step, an agent-based workflow retrieves repository functions exhibiting analogous behavior, iteratively refining the generated code by integrating semantic retrieval with feedback from conservative static analysis. Evaluated on the REPOCOD benchmark, the proposed approach achieves a Pass@1 score of 41.14%, significantly outperforming current retrieval-augmented baselines and demonstrating the efficacy of procedural similarity for repository-scale code generation.

code retrievalcross-file dependenciesprocedural similarity

This work addresses a critical gap in evaluating code-generating agents, which have predominantly focused on functional correctness while neglecting maintainability—particularly their ability to eliminate code smells and enhance long-term readability and robustness. To bridge this gap, the authors introduce SmellBench, the first systematic benchmark for assessing code refactoring capabilities through the lens of code smells. Built upon real-world open-source projects, SmellBench programmatically injects seven common code smells to create 294 high-quality, diverse refactoring tasks. The benchmark features a three-dimensional evaluation framework measuring functional correctness, accurate smell localization, and refactoring quality. Experimental results reveal that even the best-performing model combination (Qwen Code + Claude Sonnet 4.5) achieves only a score of 50.34, highlighting significant limitations in cross-file reasoning and holistic refactoring capabilities.

code agentscode smellsevaluation benchmark

Hot Scholars

MR

Michael R. Lyu

Professor of Computer Science & Engineering, The Chinese University of Hong Kong
software engineeringsoftware reliabilityfault tolerancemachine learning
DL

David Lo

Professor of Computer Science, Singapore Management University
AI4SESoftware AnalyticsSE4AISoftware Maintenance
CF

Chunrong Fang

Software Institute, Nanjing University
Software TestingSoftware EngineeringComputer Science
MP

Michael Pradel

Faculty, CISPA Helmholtz Center for Information Security • Professor, University of Stuttgart
Software EngineeringProgramming Languages
JB

Jeremiah Blocki

Purdue University
CryptographyPasswordsMemory Hard FunctionsDifferential Privacy