How Closely Do LLM Reviews Align with Human Peer Review?

๐Ÿ“… 2026-08-04
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study presents the first systematic evaluation of scientific peer review quality by three leading large language modelsโ€”GPT-5.4, Gemini 3.1 Pro Preview, and Claude Opus 4.6โ€”under a unified controlled setting, assessing their alignment with human reviewers and final ICLR 2026 decisions across 300 submissions. Through automated review generation, quantitative alignment analysis, and thematic coding, the findings reveal that all models effectively distinguish between accepted and rejected papers but fail to replicate the finer-grained distinction between oral and poster presentations. Notably, inter-model scoring biases emerge, with LLMs disproportionately emphasizing missing baselines, whereas human reviewers prioritize computational efficiency. This work provides rigorous empirical evidence on the capabilities and limitations of LLMs in scholarly peer review, offering critical insights into their reliability for academic evaluation tasks.
๐Ÿ“ Abstract
Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting. We compare reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews and final decisions for 300 topic-matched ICLR 2026 submissions, equally divided among oral, poster, and rejected papers. Each model reviewed every paper using identical instructions and rating scales after decision information was removed. Our study contributes a cross-provider analysis of three complementary dimensions: alignment with broad and fine-grained decision categories, differences in recommendation-scale usage, and thematic agreement in identified weaknesses. All three LLMs distinguished accepted from rejected papers, but none reproduced the oral versus poster distinction present in human ratings. Scoring patterns were provider-specific: Gemini assigned systematically higher ratings, while OpenAI and Claude were closer to humans for rejected and poster papers but more critical of oral papers. Human and LLM reviews also differed in emphasis, with LLMs more frequently identifying missing baseline comparisons and humans more often raising computational-efficiency concerns. These results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities.
Problem

Research questions and friction points this paper is trying to address.

LLM reviews
human peer review
decision alignment
reviewing priorities
fine-grained evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

large language models
peer review
decision alignment
review analysis
scientific evaluation
๐Ÿ”Ž Similar Papers
No similar papers found.