Score
Designs, implements, and evaluates the processes, tools, and artifacts needed to migrate large software repositories across programming languages or language versions. This includes automated source-to-source translation and refactoring, build and dependency updates, test and verification pipelines, interface and API adaptation, and performance and semantic validation to preserve functionality and maintainability at scale.
Software maintenance remains heavily reliant on manual effort, resulting in high costs, low efficiency, and susceptibility to errors. This work proposes the first systematic research framework for transfer-based software maintenance, drawing inspiration from transfer learning. The framework establishes a comprehensive lifecycle model encompassing task identification, source system selection, cross-system data matching and adaptation, and validation of transferred outcomes. It explicitly delineates the core objectives and key challenges at each stage, integrating techniques from software engineering such as knowledge transfer, cross-project data alignment, and context-aware adaptation. By doing so, the framework introduces a novel paradigm for automating software maintenance and lays a solid theoretical foundation for the future development of supporting tools and methodologies.
To address the high cost and labor-intensive nature of large-scale internal software migrations, this paper proposes the first LLM-driven, end-to-end automated migration workflow. Our method integrates a change-location identification algorithm with fine-tuned or prompt-optimized large language models (LLMs) to jointly perform migration-point detection, semantics-aware code generation, and production-ready patch synthesis—enabling a closed-loop migration process. The workflow is deeply integrated with Google’s internal codebase and CI/CD toolchain, and incorporates static code semantic analysis to ensure generation quality. Over a 12-month deployment, it successfully completed 39 migrations, producing 595 changes (93,574 edits), of which 74.45% of changes and 69.46% of edits were LLM-generated. Total migration time decreased by 50%, and developer satisfaction improved significantly. This work represents the first industrial-scale integration of LLMs across the full migration lifecycle, substantially enhancing scalability and human-AI collaboration efficiency.
Semantic version upgrades of software dependency libraries frequently break backward compatibility, necessitating automated code migration solutions. This paper proposes AIMigrate, an LLM-based migration method that innovatively incorporates version-diff information as critical contextual input to the LLM, substantially improving migration accuracy. To support this work, we construct and publicly release the first benchmark dataset specifically designed for code migration tasks, along with the end-to-end tool AIMigrate. Experimental results on real-world migration scenarios show that AIMigrate identifies 65% of necessary changes in a single inference pass and achieves 80% coverage under multiple sampling; among the generated changes, 47% are fully correct. Compared to baseline approaches using source code alone, the diff-augmented strategy consistently outperforms across multiple evaluation metrics.
This work addresses the challenge of automating library API migration in the absence of real-world migration examples. To overcome this limitation, the authors propose a novel unsupervised approach that leverages large language models (LLMs) to generate initial migration examples without requiring labeled data. These examples are then generalized by an intelligent agent into structured, testable code transformation rules, which are integrated into the PolyglotPiranha framework for execution. This study represents the first integration of LLMs’ zero-shot generation capabilities with programmatic code transformation tools. The method successfully synthesizes reusable and generalizable migration scripts across multiple Python library migration tasks, significantly enhancing the feasibility and practicality of API migration in fully unsupervised settings.
This paper addresses the challenge of ensuring functional consistency in cross-language code translation by proposing AlphaTrans, an end-to-end neural-symbolic fusion framework designed for real-world open-source projects. Methodologically, it introduces a novel *inverse call-order divide-and-conquer translation strategy*, integrating static program analysis with call-graph-driven code slicing to jointly translate source code and corresponding tests; it further establishes a three-tier automated verification mechanism—syntactic checking, unit test execution, and assertion-level output comparison. Contributions include robust support for industrial-scale projects featuring complex dependencies, custom types, and language-specific constructs. Evaluated on 10 real-world projects (836 classes, 8,575 methods, 2,719 tests), AlphaTrans achieves 96.4% syntactic correctness and 25.14% functional pass rate, with an average translation time of 34 hours per project. Developers require only 20.1 hours on average—guided by diagnostic reports—to achieve full test-suite passing.
This study addresses the failure of repository-level code migration caused by large language models overlooking cross-file dependencies. We propose a dependency-aware incremental migration framework that transcends single-file limitations by constructing dependency graphs to group translation units into dependency-consistent batches. This approach integrates compilation- and test-driven iterative verification to ensure semantic consistency. Evaluated on an industrial system comprising 51,000 lines of code, the framework achieves 100% pass rates for both compilation and testing. It significantly outperforms traditional file-level methods with faster convergence, effectively resolving critical challenges regarding completeness and scalability in large-scale code migration tasks.
为解决旧代码库迁移难题,通过ADFD-Migrate工具将程序转换为数据流图,并利用LLM生成目标语言代码,提高迁移的完整性和准确性。
This study addresses the limited understanding of how migration guides are actually provided and utilized by developers, a gap that undermines their effectiveness in managing breaking changes in software libraries. Focusing on real-world usage practices, the work presents an empirical investigation centered on libraries with incompatible updates—such as Log4j—by analyzing pull request data and patterns of documentation referencing. The findings reveal that 82.81% of references point to entire migration guides rather than specific sections, and that these guides serve not only during major version upgrades but also play a sustained role in long-term maintenance. These insights offer empirically grounded recommendations for improving the design and utility of API migration documentation.
This work addresses the challenges of IDE development posed by the rapid evolution of smart contract languages such as Move by presenting a high-performance IDE support system built atop the Move compiler and adhering to the Language Server Protocol (LSP). Through deep integration with existing language toolchains and the application of incremental parsing and optimized semantic analysis techniques, the system efficiently delivers rich IDE features even as the language undergoes continuous iteration. Deployed successfully within the Sui platform’s Move ecosystem, it significantly enhances developer experience and yields a reusable, evolution-aware IDE construction strategy applicable to other emerging programming language ecosystems.
This study addresses the challenges of high cost, error-proneness, and defect propagation in cross-repository code and test reuse during software refactoring. Through action research, the authors conduct bidirectional empirical analyses on real-world cases such as Soot/SootUp and FindBugs/SpotBugs, identifying for the first time the bidirectional reuse requirements and semantic reuse patterns inherent in refactoring scenarios. They propose a semantic alignment–based code mapping approach coupled with a hierarchical, extensible clone detection mechanism. Experimental results demonstrate that their method reduces irrelevant clones by 33%–99% on average and achieves a benchmark precision of 86%. The practical impact is further evidenced by five reported issues and ten pull requests submitted to open-source communities, eight of which have already been merged, confirming the approach’s effectiveness and applicability.