code quality

Designs, implements, and evaluates the structures, processes, and artifacts that make source code correct, readable, maintainable, and easy to change; outputs include well-tested and documented code, refactorings, linting and static-analysis rules, CI checks, and metrics to track defects and technical debt.

codequality

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

An Empirical Study on the Impact of Code Duplication-aware Refactoring Practices on Quality Metrics

Feb 06, 2025
EA
Eman Abdullah AlOmar
🏛️ Stevens Institute of Technology

This study investigates the mapping between refactoring operations performed to eliminate code duplication and software design quality metrics, along with their empirical impact. Leveraging 332 manually labeled deduplication-refactoring commits from 128 open-source Java projects, we integrate code mining, extraction of 32 structural metrics, Wilcoxon signed-rank tests, and commit semantic analysis. To our knowledge, this is the first systematic empirical validation of how widely adopted quality metrics respond to duplication-removal intent. Results show that most metrics capture this intent, yet effects are highly heterogeneous: cohesion and maintainability significantly improve, whereas complexity and coupling either remain unchanged or deteriorate. The findings expose critical limitations and contextual boundaries of conventional quality metrics in refactoring scenarios, challenging assumptions about their universality. This work provides empirical grounding for refining quality models and assessing refactoring effectiveness in practice.

Alignment of quality models with developer intentionsEvaluation of 32 structural metrics in 128 Java projectsImpact of code duplication refactoring on quality metrics

Deciphering Refactoring Branch Dynamics in Modern Code Review: An Empirical Study on Qt

Oct 07, 2024
EA
Eman Abdullah AlOmar
🏛️ Stevens Institute of Technology

This study addresses critical challenges in code review of Refactor branches within the Qt open-source project—namely, low review efficiency and insufficient documentation of developer intent. Employing a mixed-methods approach, it conducts quantitative analysis of 2,154 review records alongside manual thematic coding to construct the first comprehensive refactor-review taxonomy, comprising 12 dimensions. The analysis reveals, for the first time, that Refactor branch reviews require significantly less time yet exhibit extremely low rates of intent documentation. Based on these findings, the study derives 12 actionable, practice-oriented refactor review guidelines. Collectively, the work uncovers distinctive patterns and persistent bottlenecks in refactor review practices, thereby providing both theoretical foundations and practical guidance to enhance the quality, consistency, and traceability of refactor-related code reviews.

Code ReviewEfficiency and QualityRefactoring

This study addresses the lack of systematic evaluation of non-functional quality—specifically security, maintainability, and performance efficiency—in code generated by large language models (LLMs). Grounded in the ISO/IEC 25010 standard, it integrates a systematic literature review, dual-industry workshops, and multi-model empirical experiments (GPT-4, Claude, CodeLlama) to conduct multidimensional quality analysis on real-world software defect-fix patches. It introduces the first non-functional quality assessment framework reconciling academic rigor with industrial relevance, uncovering significant trade-offs among the three quality attributes and exposing gaps between LLM outputs and actual engineering requirements—including technical debt accumulation. Results demonstrate that functional correctness does not imply high non-functional quality, and that model architecture and optimization strategies yield markedly divergent outcomes across non-functional dimensions. The work provides both theoretical foundations and actionable guidelines for designing robust quality assurance mechanisms for LLM-generated code.

Addressing quality trade-offs in generated patches for security, maintainability and performanceEvaluating non-functional quality of LLM-generated code beyond functional correctnessInvestigating mismatches between academic focus and industry priorities on code quality

What Were You Thinking? An LLM-Driven Large-Scale Study of Refactoring Motivations in Open-Source Projects

Sep 09, 2025
MR
Mikel Robredo
🏛️ University of Oulu | University of Salerno | University of Milano-Bicocca

Prior literature inadequately characterizes developers’ real-world motivations for code refactoring in open-source projects, lacking scalable, semantically grounded analysis. Method: We introduce an LLM-driven hybrid analytical framework, performing large-scale semantic parsing of commit messages—validated via human annotation and benchmarked against traditional software metrics. Contribution/Results: Our approach uncovers 22% novel refactoring motivations absent from existing taxonomies. The LLM achieves 80% agreement with human judgments on motivation identification but aligns with established categories in only 47% of cases—demonstrating high efficacy for localized readability improvements yet revealing limitations in inferring architecture-level intent. These findings provide empirical grounding for refactoring practices and inform the design of intelligent, context-aware refactoring support tools.

Assessing software metrics correlation with motivation categoriesEvaluating LLMs in capturing refactoring motivationsIdentifying why developers refactor code

Large language models (LLMs) exhibit pervasive output formatting bias in code translation tasks—generated outputs frequently contain extraneous natural-language explanations or formatting delimiters, causing standard evaluation metrics (e.g., computation accuracy, CA) to systematically underestimate true performance. Method: We systematically evaluate 11 instruction-tuned LLMs across five programming languages and find that 26.4%–73.7% of translations require post-hoc processing to extract clean code. To address this, we propose a robust code extraction method integrating regex-based parsing with prompt engineering. Contribution/Results: Our approach achieves a 92.73% average Code Extraction Success Rate (CSR) on a multilingual alignment benchmark, substantially improving evaluation fidelity. This work is the first to quantify the impact of formatting bias and establishes a new, generalizable, and robust code extraction paradigm—providing a reproducible, standardized evaluation benchmark for LLM-based code translation.

Evaluating LLM code translation suffers from output format biasesNon-code elements in outputs interfere with performance assessment metricsProposing methods to extract source code for reliable model evaluation

Latest Papers

What's happening recently
View more

This study addresses the challenges of high cost, error-proneness, and defect propagation in cross-repository code and test reuse during software refactoring. Through action research, the authors conduct bidirectional empirical analyses on real-world cases such as Soot/SootUp and FindBugs/SpotBugs, identifying for the first time the bidirectional reuse requirements and semantic reuse patterns inherent in refactoring scenarios. They propose a semantic alignment–based code mapping approach coupled with a hierarchical, extensible clone detection mechanism. Experimental results demonstrate that their method reduces irrelevant clones by 33%–99% on average and achieves a benchmark precision of 86%. The practical impact is further evidenced by five reported issues and ten pull requests submitted to open-source communities, eight of which have already been merged, confirming the approach’s effectiveness and applicability.

clone detectioncode reusecross-repository migration

This study addresses the persistent occurrence of software defects after release, particularly in C/C++ and Java systems, whose underlying causes remain poorly understood. Through a large-scale empirical analysis of over 14,000 open-source projects, the work systematically compares pre-release and post-release defect characteristics using multidimensional metrics—including code complexity, size, change frequency, and development history—and employs statistical modeling to uncover key patterns. It reveals for the first time that post-release defects are significantly concentrated in legacy modules that undergo frequent modifications, with their root causes primarily stemming from dynamic evolutionary pressures rather than static code structure. Furthermore, such defects exhibit longer repair cycles and higher complexity, offering empirical grounding for targeted testing strategies and improved reliability assurance.

defect characterizationpost-release defectsresidual faults

This study addresses the environmental sustainability challenges arising from high energy consumption in software systems. Employing a systematic literature review, it conducts a multidimensional classification and analysis of 66 core studies. The primary contribution is the first systematic mapping framework between green code smells and refactoring techniques, precisely identifying 20 code smells that compromise environmental sustainability and proposing corresponding structured refactoring strategies. By establishing a standardized reference framework for eliminating structural inefficiencies and reducing the software ecological footprint, this work effectively advances both the theoretical foundations and practical implementation of sustainable software engineering.

Energy EfficiencyEnvironmental SustainabilityGreen Smells

Hot Scholars

DL

David Lo

Professor of Computer Science, Singapore Management University
AI4SESoftware AnalyticsSE4AISoftware Maintenance
ZP

Zifan Peng

Ph.D. Candidate at HKUST(GZ)
DeFiTrustworthy AI
IA

Iftekhar Ahmed

Associate Professor, University of California, Irvine
Software EngineeringSoftware TestingMachine Learning
FK

Foutse Khomh

NSERC Arthur B. McDonald Fellow, CRC Tier 1, Canada CIFAR AI Chair, FRQ-IVADO Chair, Full Professor
Software engineeringMachine learning systems engineeringMining software repositoriesReverse
HJ

Heng Ji

Professor of Computer Science, AICE Director, ASKS Director, UIUC, Amazon Scholar
Natural Language ProcessingLarge Language Models