ast transformations

Designs, implements, and evaluates tools or components that parse source code into abstract syntax trees (ASTs) and perform analyses, manipulations, or automated rewrites—such as tree traversals, refactorings, transpilation passes, or code-generation transformations—to produce modified ASTs, regenerated code, or analysis outputs.

asttransformations

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$181K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

AI-Driven Code Refactoring: Using Graph Neural Networks to Enhance Software Maintainability

Apr 14, 2025
GB
Gopichand Bandarupalli
🏛️ Campbellsville University

This work addresses the declining maintainability of software caused by high cyclomatic complexity and coupling in code refactoring. We propose the first end-to-end Graph Neural Network (GNN)-driven, semantics-aware refactoring method. By modeling Abstract Syntax Trees (ASTs) as graphs and integrating AST embeddings with static analysis, our approach automatically identifies and optimizes high-complexity, high-coupling code fragments. Unlike conventional rule-based or shallow-model approaches, ours is the first to systematically apply GNNs across the entire refactoring decision pipeline. Evaluated on 2 million Python code snippets, our method achieves 92% refactoring accuracy, reduces average cyclomatic complexity by 35%, and decreases coupling by 33%. These improvements significantly outperform established baselines—including SonarQube and decision tree–based methods—demonstrating both technical novelty and practical efficacy in automated, semantics-guided code refactoring.

Comparing GNN performance against rule-based and decision tree methodsReducing code complexity and coupling with AI-driven refactoringUsing GNNs to improve software maintainability via code refactoring

This study addresses the challenges of high cost, error-proneness, and defect propagation in cross-repository code and test reuse during software refactoring. Through action research, the authors conduct bidirectional empirical analyses on real-world cases such as Soot/SootUp and FindBugs/SpotBugs, identifying for the first time the bidirectional reuse requirements and semantic reuse patterns inherent in refactoring scenarios. They propose a semantic alignment–based code mapping approach coupled with a hierarchical, extensible clone detection mechanism. Experimental results demonstrate that their method reduces irrelevant clones by 33%–99% on average and achieves a benchmark precision of 86%. The practical impact is further evidenced by five reported issues and ten pull requests submitted to open-source communities, eight of which have already been merged, confirming the approach’s effectiveness and applicability.

clone detectioncode reusecross-repository migration

ASTER: Natural and Multi-language Unit Test Generation with LLMs

Sep 04, 2024
RP
Rangeet Pan
🏛️ IBM Research | Georgia Tech

To address critical challenges in multilingual (Java/Python) unit test generation—including poor test readability, low coverage, and frequent compilation failures—this paper proposes a static-analysis-driven large language model (LLM) testing paradigm. Our method integrates multilingual abstract syntax tree (AST) parsing, environment mocking, coverage-guided prompt engineering, and static analysis feedback to impose structured constraints on LLM outputs and enable iterative refinement. Evaluated on large-scale industrial Java/Python codebases and standard benchmarks, our approach achieves branch coverage competitive with or surpassing state-of-the-art (SOTA) tools; notably, it is the first to demonstrate effectiveness in Python. A user study (N=161) confirms that generated tests are significantly more natural, readable, and stylistically consistent with human-written tests.

Automated Unit TestingLow CoverageMulti-language Environment

Repository-Level Compositional Code Translation and Validation

Oct 31, 2024
AR
Ali Reza Ibrahimzada
🏛️ University of Illinois Urbana-Champaign | Indian Institute of Science | Cornell University | IBM Research

This paper addresses the challenge of ensuring functional consistency in cross-language code translation by proposing AlphaTrans, an end-to-end neural-symbolic fusion framework designed for real-world open-source projects. Methodologically, it introduces a novel *inverse call-order divide-and-conquer translation strategy*, integrating static program analysis with call-graph-driven code slicing to jointly translate source code and corresponding tests; it further establishes a three-tier automated verification mechanism—syntactic checking, unit test execution, and assertion-level output comparison. Contributions include robust support for industrial-scale projects featuring complex dependencies, custom types, and language-specific constructs. Evaluated on 10 real-world projects (836 classes, 8,575 methods, 2,719 tests), AlphaTrans achieves 96.4% syntactic correctness and 25.14% functional pass rate, with an average translation time of 34 hours per project. Developers require only 20.1 hours on average—guided by diagnostic reports—to achieve full test-suite passing.

Automate repository-level code translation across languages.Ensure functionality preservation post-translation via validation.Scale translation to real-world projects with dependencies.

Latest Papers

What's happening recently
View more

Existing direct code-to-code transformation approaches often suffer from semantic drift, implicit behavioral changes, and loss of traceability. To address these issues, this work proposes a specification-based Code2Text2Code refactoring framework that first translates source code into a neutral textual specification before generating target code. The approach integrates abstract syntax tree (AST) and dependency graph analysis, semantic-aware code chunking, retrieval-augmented generation, and DSPy-based prompt tuning, further enhanced by iterative validation and graph-based formal verification. This pipeline ensures high-fidelity semantic preservation and controllable evolution during code transformation. Experimental results demonstrate that the proposed method significantly reduces transformation loss and substantially improves semantic consistency, interface stability, and cross-language traceability of the refactored code.

behavioral changesCode2Code transformationdomain logic reconstruction

This work addresses the challenge of automating library API migration in the absence of real-world migration examples. To overcome this limitation, the authors propose a novel unsupervised approach that leverages large language models (LLMs) to generate initial migration examples without requiring labeled data. These examples are then generalized by an intelligent agent into structured, testable code transformation rules, which are integrated into the PolyglotPiranha framework for execution. This study represents the first integration of LLMs’ zero-shot generation capabilities with programmatic code transformation tools. The method successfully synthesizes reusable and generalizable migration scripts across multiple Python library migration tasks, significantly enhancing the feasibility and practicality of API migration in fully unsupervised settings.

API migrationautomated code transformationcode refactoring

Towards Cumulative Abstract Semantics via Handlers

Dec 11, 2025
CL
Cade Lueker
🏛️ University of Colorado Boulder

Modular control-flow handling in abstract interpretation and supporting multiple analysis strategies—such as path- vs. flow-sensitivity, forward vs. backward directionality, and upper vs. lower approximations—traditionally relies on complex monad transformers, leading to implementation brittleness and poor composability. Method: This paper introduces the *cumulative abstract semantics* framework, the first to incorporate *scoped effects* into abstract interpretation. It decouples syntactic structure from semantic behavior via two classes of effect handlers: *syntax-resolving* and *domain-semantics-introducing*. A single syntax-driven interpreter suffices to generate diverse dynamic evaluators and static analyzers. Contribution/Results: The framework eliminates heavyweight data structures, preserving expressiveness while drastically reducing implementation complexity for multi-strategy analyses. It enhances maintainability, composability, and modularity—providing a concise, unified, and extensible theoretical and practical foundation for modular program analysis.

Modularizing control flow in abstract interpretation frameworks.Separating syntax and semantics for flexible path and flow sensitivities.Using effects to design clean, modular interpreters and analyses.

This study addresses the unclear human-AI collaboration mechanisms in specification-driven software development with large language models (LLMs). We propose CURRANTE, a structured three-stage collaborative paradigm that guides developers through sequential refinement of requirements specifications, test cases, and function implementations. Implemented as a Visual Studio Code extension, CURRANTE integrates LLM assistance, fine-grained interaction logging, and automated test-based evaluation. By collecting interaction data and multidimensional performance metrics—including pass rates and completion time—on medium-difficulty tasks from LiveCodeBench, our work provides the first systematic empirical analysis of how iterative specification and testing dynamically influence LLM-generated code quality. These findings offer evidence-based insights for designing effective AI-augmented programming environments.

code generation qualityempirical studyhuman-in-the-loop

This study addresses the imbalance in the test pyramid—characterized by an overreliance on coarse-grained integration and system tests, which leads to difficulties in fault localization and slow execution—by proposing, for the first time, a method to automatically generate unit tests from existing integration tests. The approach combines static and dynamic analysis to automatically isolate component dependencies and enhance coverage at the unit level. Implemented as a Node.js tool and evaluated on twelve open-source JavaScript projects, the technique produces high-quality unit tests that significantly improve test suite structure, thereby increasing both testing efficiency and maintainability.

fault localizationintegration testtest pyramid

Hot Scholars

DL

David Lo

Professor of Computer Science, Singapore Management University
AI4SESoftware AnalyticsSE4AISoftware Maintenance
PL

Peng Liang

School of Computer Science, Wuhan University
Software EngineeringSoftware ArchitectureEmpirical Software Engineering
XM

Xiaoxing Ma

Professor of Computer Science and Technology, Nanjing University
software engineeringself-adaptive systemsreliability of machine learning
MC

Marcin Copik

ETH Zürich
High-Performance ComputingServerless ComputingPerformance Modeling
TH

Torsten Hoefler

Professor of Computer Science at ETH Zurich
High Performance ComputingDeep LearningNetworkingMessage Passing Interface