A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of AI reviewers to phrasing variations, which causes scoring fluctuations and erroneously rewards rhetorical optimization over scientific improvement. To this end, it formally defines “rhetorical robustness” as an independent evaluation objective and reveals the phenomenon of “spurious robustness.” The authors construct the RobustReview benchmark and propose SciCore, a dual-branch model that enhances evaluation stability through controlled full-text rewriting, structured scientific core extraction, and a dual-branch averaging fusion strategy. Empirical results demonstrate that SciCore achieves state-of-the-art joint performance in stability and discriminability while preserving alignment with human judgments, effectively mitigating sensitivity to rhetorical variations.
📝 Abstract
AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentially rewarding rhetorical optimization over scientific improvement. We formulate Rhetorical Robustness as the joint requirement of stability across content-preserving rewrites and discrimination across papers. We introduce RobustReview, a controlled full-manuscript benchmark with 1,260 manuscript versions, and evaluate 30 reviewer configurations. The benchmark reveals false robustness, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers differently. Moreover, the evaluated content-focused prompting protocol does not consistently improve robustness across backbones. Motivated by these findings, we introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a judgment based on an extracted, structured science core. This design combines manuscript-level assessment with a content-normalized view intended to reduce rhetorical sensitivity. In our primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile among the benchmarked reviewers while maintaining competitive human alignment. These results identify rhetorical robustness as a distinct evaluation target and demonstrate the potential of science-core review to improve it.
Problem

Research questions and friction points this paper is trying to address.

Trustworthy AI Reviewers
Rhetorical Robustness
Scientific Peer Review
Benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rhetorical Robustness
SciCore
RobustReview
Dual-branch Reviewer
Science Core Extraction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chenguang Wang
Virginia Tech
M
Ming Li
University of Maryland
C
Chengrui Fan
University of Maryland
Jianpeng Chen
Jianpeng Chen
Virginia Tech
Machine LearningGraph MiningMulti-view LearningAI for Science
H
Han Chen
MBZUAI
T
Tianyi Zhou
MBZUAI
Dawei Zhou
Dawei Zhou
Assistant Professor, Computer Science Department, Virginia Tech
Open-World MLAI for ScienceAI for Finance