Score
Designs, implements, and evaluates peer review systems, workflows, and protocols (including blind-review procedures), conducts technical assessments of manuscripts and grant proposals, and manages the processes and decision-making needed to navigate peer-reviewed publishing and proposal evaluation.
This study systematically investigates the application of artificial intelligence—particularly large language models—across the entire academic peer review pipeline, encompassing key stages such as review generation, author rebuttal, meta-reviewing, and manuscript revision. The work introduces the first end-to-end technical framework that integrates fine-tuning, agent-based systems, reinforcement learning, and multidimensional automated evaluation methods, comprehensively mapping existing datasets, modeling paradigms, and assessment strategies. Beyond offering practical guidelines for building end-to-end peer review assistance systems, the research also critically examines associated ethical challenges and outlines promising future directions, thereby providing both theoretical grounding and actionable insights for advancing AI-supported scholarly peer review.
This work addresses the growing strain on the peer review system in software engineering, driven by a surge in paper submissions and a scarcity of qualified reviewers. To overcome the limitations of traditional models that rely heavily on a finite pool of expert reviewers, the study proposes a novel, sustainable tripartite paradigm integrating reviewer training, community-driven incentives, and judicious AI assistance. By establishing a structured reviewer development program, implementing incentive mechanisms to foster broad community participation, and carefully deploying AI tools to enhance both efficiency and quality, the authors construct a scalable, inclusive, and resilient peer review framework. This approach aims to substantially alleviate reviewer burden, improve review quality, and encourage broader engagement from the research community.
Research Software Engineers (RSEs) face distinct challenges in ensuring code quality and maintainability, yet peer code review practices remain underexplored and inadequately supported for this community. Method: Addressing this gap, we designed and deployed a customized survey (N=61) grounded in a comparative analytical framework to systematically examine RSEs’ current review practices, key barriers, and improvement opportunities—focusing on motivations, process adaptability, tooling support, and cross-disciplinary collaboration. Contribution/Results: We identify three critical enablers of effective review adoption: lightweight process integration, domain-aware tooling, and RSE-specific training. Building on these insights, we propose a structured, ecology-oriented code review optimization framework tailored to research software. Empirical findings confirm that well-supported peer review significantly enhances the sustainability of research software, providing evidence-based guidance for RSE practice and policy development.
This study investigates key determinants of paper acceptance in open peer review, moving beyond static manuscript features to model the dynamic review process. Leveraging complete interaction data from over 28,000 submissions to ICLR 2017–2025, we integrate textual features, temporal reviewer–author behaviors, inter-reviewer disagreement metrics, and meta-review trajectories. Our analysis reveals—first time systematically—that response timeliness and interaction quality during the rebuttal phase exert a stronger influence on final decisions than initial review scores. Key findings indicate that clear writing, balanced figure usage, early submission, constructive and proactive author responses, and effective reconciliation of reviewer disagreements significantly increase acceptance probability. These results provide empirically validated, data-driven insights to enhance transparency, fairness, and efficiency in open peer review systems.
In scientific peer review, authors’ ability and effort are unobservable to journals, which must assess manuscript quality based on noisy signals—leading to suboptimal acceptance decisions. This paper proposes a dynamic review mechanism that permits authors to appeal initial rejection decisions, transforming the static, one-way review process into a two-way strategic interaction. Using game-theoretic modeling and signal-design theory, we formalize type identification, effort incentives, and optimal processing of noisy journal signals. We prove that this mechanism mitigates information asymmetry, increases the acceptance probability of high-quality manuscripts, and drives resource allocation closer to the first-best equilibrium. The key contribution lies in endogenizing the right to appeal as an incentive-compatible institutional design—marking the first such formulation in the literature—and thereby significantly improving both review efficiency and fairness.
In conference peer review, misaligned incentives among authors, conferences, and reviewers stem from inherent noise in paper quality assessment. Method: We formulate a Stackelberg game between authors and conferences, introducing the novel concept of “resubmission gap,” and analyze the dynamic trade-offs among acceptance thresholds, author resubmission behavior, and reviewer load via agent-based simulation, noise-aware modeling, and parameter estimation from historical data. Contributions/Results: (1) Raising the acceptance threshold reduces reviewer burden while preserving accepted paper quality; (2) Author self-selection induces a counterintuitive effect: stricter review increases the acceptance rate of high-quality papers; (3) A small number of high-quality reviews combined with a high threshold outperforms a large volume of low-quality reviews, and reusing prior reviews significantly alleviates load without compromising quality. We quantify how key parameters affect system performance, providing both theoretical foundations and empirical support for conference policy design.
研究通过采访技术作家探讨了软件文档审查过程及其挑战,识别了五个审查阶段,并揭示了组织和技术上的难题。
This study addresses the fragmentation of evaluation criteria for automated research systems and the difficulty of direct cross-task comparison. Employing a systematic literature review, it comprehensively examines evaluation designs across six task categories, including literature synthesis and ideation. By comparing benchmark construction and scoring protocols, this work proposes a complementary evaluation framework encompassing output-level, process-level, and human-subject assessments. It reveals the capability differences reflected by distinct designs and underscores the critical role of calibration specificity and resource budgets in performance interpretation. Furthermore, the project identifies gaps in diagnostic evaluation and provides recommendations for standardized reporting and auditing. Ultimately, these contributions offer practical guidance for benchmark selection and future research design in evaluating automated scientific discovery systems.
Traditional peer review relies excessively on authors’ narrative accounts, making it difficult to verify the authenticity of reported results. This work proposes a “code-first” review paradigm in which authors submit executable research artifacts alongside a checklist of claims. An AI-driven review infrastructure automatically provisions execution environments, runs experiments, audits code paths, and precisely maps each claim to empirical evidence, producing a standardized review package for human evaluation. By introducing AI as a core component of the review process, this study pioneers the concepts of claim-evidence contracts, generative review views, and the review package abstraction. It shifts the focus of peer review from narrative persuasion to verifiable, reproducible evidence and presents a comprehensive protocol framework encompassing system architecture, empirical validation, and analysis of governance challenges such as AI bias and prompt injection.
This study systematically investigates the fairness and effectiveness of strategyproof reviewer assignment algorithms under a two-stage peer review mechanism, specifically tailored to academic conference settings. Leveraging concepts from mechanism design in game theory, theoretical analysis, and multivariate Monte Carlo simulations, the paper evaluates the Partition and ExactDollarPartition mechanisms across varying levels of review noise, acceptance rates, reviewer loads, and inter-reviewer correlations. The findings reveal that low-noise environments disproportionately benefit marginal submissions, whereas high-noise conditions favor top-ranked ones, with mechanism performance exhibiting high sensitivity to parameter choices. This work provides the first quantitative assessment of the collective impact of strategyproof mechanisms in a two-stage review framework, offering both theoretical foundations and practical cautions for designing conference peer review systems.
Peer review is critical in software engineering research, yet systematic reviewer training has long been lacking. This work proposes and implements a large-scale Shadow Program Committee (Shadow PC) training initiative for ICSE 2026, featuring multi-stage deliberate practice, structured calibration exercises, and peer feedback mechanisms, all while maintaining strict separation from the main Program Committee to ensure review independence. Innovatively, the initiative introduces Shadow PC Area Chairs to establish a sustainable leadership development pipeline. With 102 participants completing reviews for 117 papers, 97% of participants recommended the experience, and 67% of authors found the reviews helpful, demonstrating the feasibility of scaling high-quality reviewer training effectively.