🤖 AI Summary
This study addresses the lack of standardized guidelines and systematic evaluation of AI-assisted peer review in academic publishing. It presents the first interdisciplinary analysis of AI usage policies among reviewers across 111 AI/NLP conferences and medical journals, and introduces a multidimensional framework for assessing review quality based on newly collected datasets from ICLR 2026 and Nature Communications. Leveraging methods such as LLM-as-a-Judge, scoring consistency, fine-grained content analysis, and overlap with human judgments, the work compares the performance of open- and closed-source large language models. Findings reveal that while AI systems produce fluent and detailed reviews, they exhibit systematic shortcomings—including overly optimistic recommendations, overly generalized critiques, and insufficient evidential support—and that aggregate scores alone tend to overestimate their actual review quality.
📝 Abstract
AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities. Second, we evaluate AI-generated peer reviews at ICLR 2026 and Nature Communications using a novel dataset comprising original manuscript submissions and several hundred human- and machine-generated reviews. We compare reviews produced by open-source and proprietary models using complementary evaluation metrics, including LLM-as-a-Judge, score alignment, granularity, and overlap with human reviewers' concerns. Our results show that current LLMs can generate detailed and fluent reviews but exhibit systematic weaknesses, such as overly positive recommendations, generic criticism, and uneven evidence grounding. We demonstrate that aggregate quality scores alone can overestimate review quality and argue for multi-dimensional evaluation of AI-generated peer reviews.