Score
Designs, implements, and evaluates the structures, processes, and artifacts that make source code correct, readable, maintainable, and easy to change; outputs include well-tested and documented code, refactorings, linting and static-analysis rules, CI checks, and metrics to track defects and technical debt.
This study investigates the mapping between refactoring operations performed to eliminate code duplication and software design quality metrics, along with their empirical impact. Leveraging 332 manually labeled deduplication-refactoring commits from 128 open-source Java projects, we integrate code mining, extraction of 32 structural metrics, Wilcoxon signed-rank tests, and commit semantic analysis. To our knowledge, this is the first systematic empirical validation of how widely adopted quality metrics respond to duplication-removal intent. Results show that most metrics capture this intent, yet effects are highly heterogeneous: cohesion and maintainability significantly improve, whereas complexity and coupling either remain unchanged or deteriorate. The findings expose critical limitations and contextual boundaries of conventional quality metrics in refactoring scenarios, challenging assumptions about their universality. This work provides empirical grounding for refining quality models and assessing refactoring effectiveness in practice.
This study addresses critical challenges in code review of Refactor branches within the Qt open-source project—namely, low review efficiency and insufficient documentation of developer intent. Employing a mixed-methods approach, it conducts quantitative analysis of 2,154 review records alongside manual thematic coding to construct the first comprehensive refactor-review taxonomy, comprising 12 dimensions. The analysis reveals, for the first time, that Refactor branch reviews require significantly less time yet exhibit extremely low rates of intent documentation. Based on these findings, the study derives 12 actionable, practice-oriented refactor review guidelines. Collectively, the work uncovers distinctive patterns and persistent bottlenecks in refactor review practices, thereby providing both theoretical foundations and practical guidance to enhance the quality, consistency, and traceability of refactor-related code reviews.
This study addresses the lack of systematic evaluation of non-functional quality—specifically security, maintainability, and performance efficiency—in code generated by large language models (LLMs). Grounded in the ISO/IEC 25010 standard, it integrates a systematic literature review, dual-industry workshops, and multi-model empirical experiments (GPT-4, Claude, CodeLlama) to conduct multidimensional quality analysis on real-world software defect-fix patches. It introduces the first non-functional quality assessment framework reconciling academic rigor with industrial relevance, uncovering significant trade-offs among the three quality attributes and exposing gaps between LLM outputs and actual engineering requirements—including technical debt accumulation. Results demonstrate that functional correctness does not imply high non-functional quality, and that model architecture and optimization strategies yield markedly divergent outcomes across non-functional dimensions. The work provides both theoretical foundations and actionable guidelines for designing robust quality assurance mechanisms for LLM-generated code.
Prior literature inadequately characterizes developers’ real-world motivations for code refactoring in open-source projects, lacking scalable, semantically grounded analysis. Method: We introduce an LLM-driven hybrid analytical framework, performing large-scale semantic parsing of commit messages—validated via human annotation and benchmarked against traditional software metrics. Contribution/Results: Our approach uncovers 22% novel refactoring motivations absent from existing taxonomies. The LLM achieves 80% agreement with human judgments on motivation identification but aligns with established categories in only 47% of cases—demonstrating high efficacy for localized readability improvements yet revealing limitations in inferring architecture-level intent. These findings provide empirical grounding for refactoring practices and inform the design of intelligent, context-aware refactoring support tools.
Large language models (LLMs) exhibit pervasive output formatting bias in code translation tasks—generated outputs frequently contain extraneous natural-language explanations or formatting delimiters, causing standard evaluation metrics (e.g., computation accuracy, CA) to systematically underestimate true performance. Method: We systematically evaluate 11 instruction-tuned LLMs across five programming languages and find that 26.4%–73.7% of translations require post-hoc processing to extract clean code. To address this, we propose a robust code extraction method integrating regex-based parsing with prompt engineering. Contribution/Results: Our approach achieves a 92.73% average Code Extraction Success Rate (CSR) on a multilingual alignment benchmark, substantially improving evaluation fidelity. This work is the first to quantify the impact of formatting bias and establishes a new, generalizable, and robust code extraction paradigm—providing a reproducible, standardized evaluation benchmark for LLM-based code translation.
为了解决程序修复过程中需求到修复过程的可审查性问题,提出THEMIS方法,通过语义解析、需求-代码图等手段实现过程外部化。
本文提出一种生命周期感知的框架,结合软件质量评估与大语言模型代码优化,以解决科研软件因开发者缺乏软件工程经验导致的质量问题。
This study addresses the challenges of high cost, error-proneness, and defect propagation in cross-repository code and test reuse during software refactoring. Through action research, the authors conduct bidirectional empirical analyses on real-world cases such as Soot/SootUp and FindBugs/SpotBugs, identifying for the first time the bidirectional reuse requirements and semantic reuse patterns inherent in refactoring scenarios. They propose a semantic alignment–based code mapping approach coupled with a hierarchical, extensible clone detection mechanism. Experimental results demonstrate that their method reduces irrelevant clones by 33%–99% on average and achieves a benchmark precision of 86%. The practical impact is further evidenced by five reported issues and ten pull requests submitted to open-source communities, eight of which have already been merged, confirming the approach’s effectiveness and applicability.
This study addresses the persistent occurrence of software defects after release, particularly in C/C++ and Java systems, whose underlying causes remain poorly understood. Through a large-scale empirical analysis of over 14,000 open-source projects, the work systematically compares pre-release and post-release defect characteristics using multidimensional metrics—including code complexity, size, change frequency, and development history—and employs statistical modeling to uncover key patterns. It reveals for the first time that post-release defects are significantly concentrated in legacy modules that undergo frequent modifications, with their root causes primarily stemming from dynamic evolutionary pressures rather than static code structure. Furthermore, such defects exhibit longer repair cycles and higher complexity, offering empirical grounding for targeted testing strategies and improved reliability assurance.
This study addresses the environmental sustainability challenges arising from high energy consumption in software systems. Employing a systematic literature review, it conducts a multidimensional classification and analysis of 66 core studies. The primary contribution is the first systematic mapping framework between green code smells and refactoring techniques, precisely identifying 20 code smells that compromise environmental sustainability and proposing corresponding structured refactoring strategies. By establishing a standardized reference framework for eliminating structural inefficiencies and reducing the software ecological footprint, this work effectively advances both the theoretical foundations and practical implementation of sustainable software engineering.