code refactoring

Designs and implements code transformations and migration plans that restructure source code, modules, components, frameworks, and system architecture to improve maintainability, modularity, performance, or adaptability, including incremental and large‑scale refactorings of legacy codebases. Builds and evaluates refactoring strategies, patterns, practices, and automated refactoring tools or scripts to reliably apply, verify, and analyze refactorings across codebases.

coderefactoring

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.73
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$191K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of a scalable, traceable, and systematic approach to modernizing large-scale legacy systems while preserving both functional and non-functional characteristics. The authors propose a four-phase model-driven method that leverages a semantically rich intermediate model to uniformly abstract a legacy system’s structure, dependencies, and metadata. By designing semantics-preserving transformation rules, the approach enables semi-automated migration to modern platforms such as web-based architectures. The method establishes an end-to-end model-driven pipeline that integrates semantic metadata modeling with automated code synthesis. Evaluated on an industrial-scale .NET system, it successfully migrated core UI components, significantly enhancing maintainability and scalability while reducing modernization risks and manual effort.

intermediate modellegacy system modernizationmodel-driven engineering

This work addresses the risk that automated Python refactoring tools may inadvertently introduce behavioral changes, thereby compromising software reliability. To tackle this issue, the authors propose a novel approach that leverages foundation models as semantic oracles, integrated with Git diff parsing and automated validation, to detect behavior-altering refactorings. Applying this method to 217 refactoring instances produced by the Rope tool, the study uncovers 13 previously unknown defects, 12 of which have been acknowledged and fixed by the developers. This demonstrates the effectiveness of the technique in enhancing the trustworthiness and practical utility of automated refactoring tools.

automated code transformationbehavioral changesPython refactoring

Refactoring Towards Microservices: Preparing the Ground for Service Extraction

Oct 03, 2025
RP
Rita Peixoto
🏛️ INESC TEC | Faculty of Engineering | University of Porto | IFTO | UNIBZ | IME/USP

Migrating monolithic systems to microservices faces a critical challenge: the lack of systematic, code-level guidance for identifying and decoupling inter-component dependencies—existing research predominantly addresses architectural concerns while neglecting actionable, refactor-driven practices. To bridge this gap, we propose a code-level refactoring methodology tailored for microservice migration. Our approach introduces the first comprehensive refactoring catalog for migration, comprising seven empirically grounded patterns that address key scenarios—including service boundary identification, cross-service call extraction, and data decoupling. Integrating literature analysis with industrial practice, the method leverages dependency graph analysis, semantics-aware refactoring, and a hierarchical classification strategy to enable standardized and automatable migration. Experimental evaluation demonstrates that our approach significantly reduces refactoring decision complexity, improves service extraction accuracy and long-term maintainability, and delivers the first production-ready, extensible code-level migration framework for microservice evolution.

Addressing code-level challenges in monolithic to microservices migrationProviding systematic refactorings to handle service dependencies effectivelySimplifying manual migration process through structured step-by-step approach

This work addresses the challenge of automating library API migration in the absence of real-world migration examples. To overcome this limitation, the authors propose a novel unsupervised approach that leverages large language models (LLMs) to generate initial migration examples without requiring labeled data. These examples are then generalized by an intelligent agent into structured, testable code transformation rules, which are integrated into the PolyglotPiranha framework for execution. This study represents the first integration of LLMs’ zero-shot generation capabilities with programmatic code transformation tools. The method successfully synthesizes reusable and generalizable migration scripts across multiple Python library migration tasks, significantly enhancing the feasibility and practicality of API migration in fully unsupervised settings.

API migrationautomated code transformationcode refactoring

Latest Papers

What's happening recently
View more

This study addresses the challenges of high cost, error-proneness, and defect propagation in cross-repository code and test reuse during software refactoring. Through action research, the authors conduct bidirectional empirical analyses on real-world cases such as Soot/SootUp and FindBugs/SpotBugs, identifying for the first time the bidirectional reuse requirements and semantic reuse patterns inherent in refactoring scenarios. They propose a semantic alignment–based code mapping approach coupled with a hierarchical, extensible clone detection mechanism. Experimental results demonstrate that their method reduces irrelevant clones by 33%–99% on average and achieves a benchmark precision of 86%. The practical impact is further evidenced by five reported issues and ten pull requests submitted to open-source communities, eight of which have already been merged, confirming the approach’s effectiveness and applicability.

clone detectioncode reusecross-repository migration

Software maintenance remains heavily reliant on manual effort, resulting in high costs, low efficiency, and susceptibility to errors. This work proposes the first systematic research framework for transfer-based software maintenance, drawing inspiration from transfer learning. The framework establishes a comprehensive lifecycle model encompassing task identification, source system selection, cross-system data matching and adaptation, and validation of transferred outcomes. It explicitly delineates the core objectives and key challenges at each stage, integrating techniques from software engineering such as knowledge transfer, cross-project data alignment, and context-aware adaptation. By doing so, the framework introduces a novel paradigm for automating software maintenance and lays a solid theoretical foundation for the future development of supporting tools and methodologies.

automated maintenanceknowledge transfermigration-based maintenance

This work addresses the challenge that code generated by large language models (LLMs) often suffers from high complexity, redundancy, and architectural debt, and struggles to autonomously identify and perform human-level refactoring. To this end, we introduce CodeTaste, a benchmark that systematically evaluates LLMs’ ability to detect and reproduce real-world refactorings in multi-file settings. Our approach combines large-scale mining of open-source changes, data-flow analysis, and static pattern detection, employing a two-stage “propose-and-implement” strategy. Refactoring quality is validated through test suites and behavioral equivalence checks. Experiments show that while current LLMs can effectively refactor under explicit instructions, they still exhibit a significant gap in autonomously understanding human refactoring intent. Performance is notably enhanced by the staged strategy and by prioritizing proposals aligned with human practices.

behavior-preserving transformationcode qualitycode refactoring

Large language models (LLMs) have gained widespread popularity and have steadily improved over time, enabling software developers to use them for various code-related tasks. One common task is code refactoring, where the LLM suggests changes for the developer to apply to their code to improve quality attributes such as readability or maintainability. While current research focuses on evaluating LLM-generated refactoring suggestions, there is a limited understanding of how developers apply these suggestions in practice. To explore this, we analyze 169 GitHub commits where developers refactor their code based on a ChatGPT conversation linked in the commit message. We found that developers mostly accept and use the suggestions without modifications. When changes are made, they are mostly major and fall into five different patterns that depend on the refactoring activity, the developer's prompt, and the validity of the response from ChatGPT.

ChatGPTcode refactoringdeveloper adoption

Existing direct code-to-code transformation approaches often suffer from semantic drift, implicit behavioral changes, and loss of traceability. To address these issues, this work proposes a specification-based Code2Text2Code refactoring framework that first translates source code into a neutral textual specification before generating target code. The approach integrates abstract syntax tree (AST) and dependency graph analysis, semantic-aware code chunking, retrieval-augmented generation, and DSPy-based prompt tuning, further enhanced by iterative validation and graph-based formal verification. This pipeline ensures high-fidelity semantic preservation and controllable evolution during code transformation. Experimental results demonstrate that the proposed method significantly reduces transformation loss and substantially improves semantic consistency, interface stability, and cross-language traceability of the refactored code.

behavioral changesCode2Code transformationdomain logic reconstruction

Hot Scholars

TN

Tien N. Nguyen

Professor, School of Engineering and Computer Science - The University of Texas at Dallas
AI4SEAutomated Software EngineeringArtificial IntelligenceMining Software Repositories
SM

Stephen M. Watt

University of Waterloo
computer algebracompilershandwriting recognition
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
MP

Michael Pradel

Faculty, CISPA Helmholtz Center for Information Security • Professor, University of Stuttgart
Software EngineeringProgramming Languages
AG

Alessandro Garcia

Associate Professor, Computer Science, Pontifical Catholic University of Rio de Janeiro
Software Engineering