Score
Performs systematic examination of source code, design documents, and change sets to identify defects, architectural inconsistencies, security and performance issues, and violations of style or standards. Produces actionable feedback such as annotated comments, suggested fixes, review reports, and acceptance or rejection decisions to drive corrective changes and maintain code and design quality.
This study addresses critical challenges in code review of Refactor branches within the Qt open-source project—namely, low review efficiency and insufficient documentation of developer intent. Employing a mixed-methods approach, it conducts quantitative analysis of 2,154 review records alongside manual thematic coding to construct the first comprehensive refactor-review taxonomy, comprising 12 dimensions. The analysis reveals, for the first time, that Refactor branch reviews require significantly less time yet exhibit extremely low rates of intent documentation. Based on these findings, the study derives 12 actionable, practice-oriented refactor review guidelines. Collectively, the work uncovers distinctive patterns and persistent bottlenecks in refactor review practices, thereby providing both theoretical foundations and practical guidance to enhance the quality, consistency, and traceability of refactor-related code reviews.
This work addresses the lack of a well-defined, actionable metric for assessing the clarity of code review comments (CRCs). We propose RIE—the first industrially grounded, three-dimensional clarity model comprising Relevance, Informativeness, and Expressiveness—and introduce ClearCRC, an automated evaluation framework. Through a systematic literature review, a developer survey (N=217), and empirical analysis across nine programming languages in open-source projects, we find that 28.8% of CRCs exhibit clarity deficiencies in at least one RIE dimension. ClearCRC integrates rule-based and heuristic features, significantly outperforming baseline methods in accuracy and F1-score. This study is the first to systematically characterize developers’ real-world expectations regarding CRC clarity, thereby establishing both a theoretical foundation and a practical toolset for improving code review quality.
This study addresses the time-consuming and inefficient nature of manual code review by conducting a structured systematic literature review (SLR) of 119 publications—the first to propose a comprehensive, task-dimensional taxonomy for automated code review. Methodologically, it integrates machine learning, information retrieval, program analysis, and natural language processing techniques, and empirically evaluates approaches using datasets from GitHub, Gerrit, and other platforms, with metrics including BLEU, F1, and MAP. Key contributions are: (1) a refined classification of 12 automated review tasks and 7 core technical paradigms; (2) a curated inventory of 32 publicly available tools and datasets; (3) identification of critical bottlenecks in data-driven methods—particularly regarding interpretability and cross-project generalizability; and (4) a reproducible evaluation benchmark alongside four concrete directions for future research. The findings are synthesized into a rigorous, structured SLR report.
This study investigates which types of code review comments developers are more likely to adopt, aiming to guide LLMs in generating high-adoption-rate review suggestions. We propose a five-dimensional taxonomy—readability, defects, maintainability, design, and style—and develop an LLM-as-a-Judge framework to automatically classify both human-written and LLM-generated review comments from real-world projects. Empirical analysis reveals that comments targeting readability, defects, and maintainability achieve significantly higher resolution rates; LLM-generated comments exhibit strong actionability overall and demonstrate complementary strengths relative to human reviewers across diverse project contexts. Our key contributions are: (1) the first systematic empirical characterization of the relationship between comment type and developer adoption behavior, and (2) validation of the LLM-as-a-Judge paradigm for code review quality assessment, demonstrating its effectiveness and generalizability across projects and comment categories.
This work proposes a fine-tuning-free, large language model (LLM)-driven approach to address the need for high-quality, context-aware, and goal-directed automated code review comments in enterprise settings. By leveraging prompt engineering, contextual retrieval, and a comment quality filtering mechanism, the authors developed and deployed RovoDev Code Reviewer—an integrated system within Atlassian Bitbucket. Evaluation over a one-year period in a real-world industrial environment demonstrates that 38.7% of the system’s automatically generated comments led developers to modify their code, resulting in a 30.8% reduction in average pull request (PR) cycle time and a 35.6% decrease in manual reviewer comments. The system also effectively identified actionable code defects, confirming its practicality and effectiveness without requiring model fine-tuning.
This study addresses the limited understanding of how developers respond to code review comments generated by AI coding agents, a gap that hinders the effective integration of AI-assisted reviewing. Through a large-scale empirical analysis of 54,791 AI-generated review comments across 342 Python repositories—produced by prominent agents including Copilot, Cursor, Codex, Devin, and Claude—the authors combine GitHub data mining, quantitative statistics, and open card sorting to uncover real-world developer response patterns. The findings reveal that Copilot accounts for 72.9% of resolved comments; core developers primarily address design-related feedback, while peripheral contributors focus on functional defects; unresolved comments often stem from incorrect suggestions or deliberate design choices; and inline code suggestions significantly increase adoption rates. The study further identifies ten root causes underlying unresolved feedback.
This study addresses the lack of effective workflows and IDE tools that support end-to-end trust calibration for developers reviewing multi-file code changes generated by large language models (LLMs). In collaboration with JetBrains, the authors employed a double-diamond design process through participatory design to propose a three-tiered review workflow—comprising overview, file-level analysis, and code snippet inspection—centered on trust calibration, along with seven key design components. A high-fidelity, semi-interactive prototype was developed and evaluated, with results showing that the three-tiered workflow received significantly higher ratings than a neutral baseline. Notably, 63% of participants anticipated reduced review effort, and 52% reported a lower burden in assessing trustworthiness, demonstrating the framework’s effectiveness and potential for building AI-ready code review tools.