Score
Designs, implements, and evaluates code review systems and practices, including review workflows, checklists, templates, metrics, and automation rules; integrates reviews with CI/CD and static-analysis tools and produces sample code and review artifacts to enforce clean code principles. Coaches and mentors reviewers, sets governance and cultural norms for peer and design reviews, and analyzes process effectiveness and code-quality outcomes to improve review discipline and leadership.
This study addresses the time-consuming and inefficient nature of manual code review by conducting a structured systematic literature review (SLR) of 119 publications—the first to propose a comprehensive, task-dimensional taxonomy for automated code review. Methodologically, it integrates machine learning, information retrieval, program analysis, and natural language processing techniques, and empirically evaluates approaches using datasets from GitHub, Gerrit, and other platforms, with metrics including BLEU, F1, and MAP. Key contributions are: (1) a refined classification of 12 automated review tasks and 7 core technical paradigms; (2) a curated inventory of 32 publicly available tools and datasets; (3) identification of critical bottlenecks in data-driven methods—particularly regarding interpretability and cross-project generalizability; and (4) a reproducible evaluation benchmark alongside four concrete directions for future research. The findings are synthesized into a rigorous, structured SLR report.
Research Software Engineers (RSEs) face distinct challenges in ensuring code quality and maintainability, yet peer code review practices remain underexplored and inadequately supported for this community. Method: Addressing this gap, we designed and deployed a customized survey (N=61) grounded in a comparative analytical framework to systematically examine RSEs’ current review practices, key barriers, and improvement opportunities—focusing on motivations, process adaptability, tooling support, and cross-disciplinary collaboration. Contribution/Results: We identify three critical enablers of effective review adoption: lightweight process integration, domain-aware tooling, and RSE-specific training. Building on these insights, we propose a structured, ecology-oriented code review optimization framework tailored to research software. Empirical findings confirm that well-supported peer review significantly enhances the sustainability of research software, providing evidence-based guidance for RSE practice and policy development.
Peer review is critical in software engineering research, yet systematic reviewer training has long been lacking. This work proposes and implements a large-scale Shadow Program Committee (Shadow PC) training initiative for ICSE 2026, featuring multi-stage deliberate practice, structured calibration exercises, and peer feedback mechanisms, all while maintaining strict separation from the main Program Committee to ensure review independence. Innovatively, the initiative introduces Shadow PC Area Chairs to establish a sustainable leadership development pipeline. With 102 participants completing reviews for 117 papers, 97% of participants recommended the experience, and 67% of authors found the reviews helpful, demonstrating the feasibility of scaling high-quality reviewer training effectively.
This study investigates the cognitive mechanisms underlying experienced reviewers’ comprehension of code changes during real-world code reviews. Method: Grounded in Letovsky’s theory of program comprehension, we conducted a qualitative study employing theory-driven thematic analysis, on-site observations, and semi-structured interviews. Contribution/Results: We present the first domain-specific cognitive model for code review—Code Review Comprehension Model (CRCM)—which characterizes review comprehension as a three-stage process: *context construction*, *multimodal inspection*, and *mental model comparison*. Our findings empirically establish code comprehension as the foundational cognitive activity in effective review, identify reusable, evidence-based review strategies, and provide actionable design implications for review tooling—including context augmentation and mental model visualization—to enhance feedback quality and review efficiency. This work marks the first rigorous application of program comprehension theory to empirical code review practice.
This work proposes a fine-tuning-free, large language model (LLM)-driven approach to address the need for high-quality, context-aware, and goal-directed automated code review comments in enterprise settings. By leveraging prompt engineering, contextual retrieval, and a comment quality filtering mechanism, the authors developed and deployed RovoDev Code Reviewer—an integrated system within Atlassian Bitbucket. Evaluation over a one-year period in a real-world industrial environment demonstrates that 38.7% of the system’s automatically generated comments led developers to modify their code, resulting in a 30.8% reduction in average pull request (PR) cycle time and a 35.6% decrease in manual reviewer comments. The system also effectively identified actionable code defects, confirming its practicality and effectiveness without requiring model fine-tuning.
GitHub’s CODEOWNERS feature automates code review responsibility assignment, yet its real-world adoption and impact remain poorly understood. This paper presents the first large-scale empirical study, analyzing 840,000 pull requests (PRs) and 2 million review logs across 2,147 open-source projects using Regression Discontinuity Design (RDD) to causally quantify CODEOWNERS’ effects. Results show that CODEOWNERS significantly improves review timeliness and coverage while reducing review burden on core developers; promotes more equitable ownership distribution and accelerates PR integration; and functions as a novel software governance mechanism that enhances project security and collaborative resilience. Collectively, this work demonstrates that automated ownership assignment substantively reshapes both collaboration efficiency and governance structures in open-source development.
This study addresses the challenges of scaling code review in software engineering education—namely, time constraints, inconsistent feedback, and students’ limited experience—by integrating large language models (LLMs) into the GitHub pull request (PR) workflow. The authors propose an in-workflow human-AI collaborative review mechanism that enables students to conduct authentic code reviews and fosters self-regulated learning. Through a controlled experiment across two course offerings, combining GitHub log data with student reflections, the study finds that the 2024 cohort submitted significantly more PRs (1,176 vs. 581), achieved zero AI invocation failures, and demonstrated substantive improvements in approximately one-third of AI-reviewed PRs. These results suggest students focused more on code quality discussions and reduced overreliance on AI, offering empirical support and design insights for AI-augmented code review pedagogy.
This work addresses a critical gap in the evaluation of intelligent code review systems, which has predominantly emphasized performance metrics while overlooking agents’ dynamic behaviors, failure modes, and implicit operational costs in real-world developer environments. We propose a trajectory-aware, cost-sensitive evaluation framework that systematically analyzes authentic code review logs from local development settings by integrating trajectory parsing, behavioral modeling, and overhead quantification to assess both planning efficacy and verification costs. Our findings reveal that high-precision reviews often incur substantial exploration and validation overhead, whereas successful cases consistently exhibit stronger upfront planning capabilities that significantly reduce downstream verification burden. This study is the first to incorporate trajectory-level behavior and associated costs into the evaluation paradigm, uncovering key trade-offs essential for the practical deployment of intelligent code review agents.
研究通过采访技术作家探讨了软件文档审查过程及其挑战,识别了五个审查阶段,并揭示了组织和技术上的难题。
This work proposes a novel AI-agent-driven, end-to-end code review paradigm to address the challenges posed by the exponential growth of code generated by AI programming tools and the inefficiency and high cognitive load of traditional manual code reviews. The framework integrates large language models and multi-agent systems across five stages—pull request (PR) creation, enhancement, reviewer assignment, AI-assisted review, and PR retrospection—to enable automated processing, intelligent recommendations, and comment generation, while embedding human-in-the-loop quality gates at critical junctures to ensure accountability and transparency. This study presents the first systematic architecture for code review in the era of large models, overcoming the fragmentation of existing tools and identifying six key open challenges—including reliability, bias, and privacy—to establish a clear research agenda for human-AI collaborative software engineering.