Score
Designs, implements, and evaluates tools or components that parse source code into abstract syntax trees (ASTs) and perform analyses, manipulations, or automated rewrites—such as tree traversals, refactorings, transpilation passes, or code-generation transformations—to produce modified ASTs, regenerated code, or analysis outputs.
This work addresses the declining maintainability of software caused by high cyclomatic complexity and coupling in code refactoring. We propose the first end-to-end Graph Neural Network (GNN)-driven, semantics-aware refactoring method. By modeling Abstract Syntax Trees (ASTs) as graphs and integrating AST embeddings with static analysis, our approach automatically identifies and optimizes high-complexity, high-coupling code fragments. Unlike conventional rule-based or shallow-model approaches, ours is the first to systematically apply GNNs across the entire refactoring decision pipeline. Evaluated on 2 million Python code snippets, our method achieves 92% refactoring accuracy, reduces average cyclomatic complexity by 35%, and decreases coupling by 33%. These improvements significantly outperform established baselines—including SonarQube and decision tree–based methods—demonstrating both technical novelty and practical efficacy in automated, semantics-guided code refactoring.
This study addresses the challenges of high cost, error-proneness, and defect propagation in cross-repository code and test reuse during software refactoring. Through action research, the authors conduct bidirectional empirical analyses on real-world cases such as Soot/SootUp and FindBugs/SpotBugs, identifying for the first time the bidirectional reuse requirements and semantic reuse patterns inherent in refactoring scenarios. They propose a semantic alignment–based code mapping approach coupled with a hierarchical, extensible clone detection mechanism. Experimental results demonstrate that their method reduces irrelevant clones by 33%–99% on average and achieves a benchmark precision of 86%. The practical impact is further evidenced by five reported issues and ten pull requests submitted to open-source communities, eight of which have already been merged, confirming the approach’s effectiveness and applicability.
To address critical challenges in multilingual (Java/Python) unit test generation—including poor test readability, low coverage, and frequent compilation failures—this paper proposes a static-analysis-driven large language model (LLM) testing paradigm. Our method integrates multilingual abstract syntax tree (AST) parsing, environment mocking, coverage-guided prompt engineering, and static analysis feedback to impose structured constraints on LLM outputs and enable iterative refinement. Evaluated on large-scale industrial Java/Python codebases and standard benchmarks, our approach achieves branch coverage competitive with or surpassing state-of-the-art (SOTA) tools; notably, it is the first to demonstrate effectiveness in Python. A user study (N=161) confirms that generated tests are significantly more natural, readable, and stylistically consistent with human-written tests.
This paper addresses the challenge of ensuring functional consistency in cross-language code translation by proposing AlphaTrans, an end-to-end neural-symbolic fusion framework designed for real-world open-source projects. Methodologically, it introduces a novel *inverse call-order divide-and-conquer translation strategy*, integrating static program analysis with call-graph-driven code slicing to jointly translate source code and corresponding tests; it further establishes a three-tier automated verification mechanism—syntactic checking, unit test execution, and assertion-level output comparison. Contributions include robust support for industrial-scale projects featuring complex dependencies, custom types, and language-specific constructs. Evaluated on 10 real-world projects (836 classes, 8,575 methods, 2,719 tests), AlphaTrans achieves 96.4% syntactic correctness and 25.14% functional pass rate, with an average translation time of 34 hours per project. Developers require only 20.1 hours on average—guided by diagnostic reports—to achieve full test-suite passing.
Existing direct code-to-code transformation approaches often suffer from semantic drift, implicit behavioral changes, and loss of traceability. To address these issues, this work proposes a specification-based Code2Text2Code refactoring framework that first translates source code into a neutral textual specification before generating target code. The approach integrates abstract syntax tree (AST) and dependency graph analysis, semantic-aware code chunking, retrieval-augmented generation, and DSPy-based prompt tuning, further enhanced by iterative validation and graph-based formal verification. This pipeline ensures high-fidelity semantic preservation and controllable evolution during code transformation. Experimental results demonstrate that the proposed method significantly reduces transformation loss and substantially improves semantic consistency, interface stability, and cross-language traceability of the refactored code.
This work addresses the challenge of automating library API migration in the absence of real-world migration examples. To overcome this limitation, the authors propose a novel unsupervised approach that leverages large language models (LLMs) to generate initial migration examples without requiring labeled data. These examples are then generalized by an intelligent agent into structured, testable code transformation rules, which are integrated into the PolyglotPiranha framework for execution. This study represents the first integration of LLMs’ zero-shot generation capabilities with programmatic code transformation tools. The method successfully synthesizes reusable and generalizable migration scripts across multiple Python library migration tasks, significantly enhancing the feasibility and practicality of API migration in fully unsupervised settings.
Modular control-flow handling in abstract interpretation and supporting multiple analysis strategies—such as path- vs. flow-sensitivity, forward vs. backward directionality, and upper vs. lower approximations—traditionally relies on complex monad transformers, leading to implementation brittleness and poor composability. Method: This paper introduces the *cumulative abstract semantics* framework, the first to incorporate *scoped effects* into abstract interpretation. It decouples syntactic structure from semantic behavior via two classes of effect handlers: *syntax-resolving* and *domain-semantics-introducing*. A single syntax-driven interpreter suffices to generate diverse dynamic evaluators and static analyzers. Contribution/Results: The framework eliminates heavyweight data structures, preserving expressiveness while drastically reducing implementation complexity for multi-strategy analyses. It enhances maintainability, composability, and modularity—providing a concise, unified, and extensible theoretical and practical foundation for modular program analysis.
This study addresses the unclear human-AI collaboration mechanisms in specification-driven software development with large language models (LLMs). We propose CURRANTE, a structured three-stage collaborative paradigm that guides developers through sequential refinement of requirements specifications, test cases, and function implementations. Implemented as a Visual Studio Code extension, CURRANTE integrates LLM assistance, fine-grained interaction logging, and automated test-based evaluation. By collecting interaction data and multidimensional performance metrics—including pass rates and completion time—on medium-difficulty tasks from LiveCodeBench, our work provides the first systematic empirical analysis of how iterative specification and testing dynamically influence LLM-generated code quality. These findings offer evidence-based insights for designing effective AI-augmented programming environments.
This study addresses the imbalance in the test pyramid—characterized by an overreliance on coarse-grained integration and system tests, which leads to difficulties in fault localization and slow execution—by proposing, for the first time, a method to automatically generate unit tests from existing integration tests. The approach combines static and dynamic analysis to automatically isolate component dependencies and enhance coverage at the unit level. Implemented as a Node.js tool and evaluated on twelve open-source JavaScript projects, the technique produces high-quality unit tests that significantly improve test suite structure, thereby increasing both testing efficiency and maintainability.