🤖 AI Summary
This study addresses the vulnerability of AI reviewers to phrasing variations, which causes scoring fluctuations and erroneously rewards rhetorical optimization over scientific improvement. To this end, it formally defines “rhetorical robustness” as an independent evaluation objective and reveals the phenomenon of “spurious robustness.” The authors construct the RobustReview benchmark and propose SciCore, a dual-branch model that enhances evaluation stability through controlled full-text rewriting, structured scientific core extraction, and a dual-branch averaging fusion strategy. Empirical results demonstrate that SciCore achieves state-of-the-art joint performance in stability and discriminability while preserving alignment with human judgments, effectively mitigating sensitivity to rhetorical variations.
📝 Abstract
AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentially rewarding rhetorical optimization over scientific improvement. We formulate Rhetorical Robustness as the joint requirement of stability across content-preserving rewrites and discrimination across papers. We introduce RobustReview, a controlled full-manuscript benchmark with 1,260 manuscript versions, and evaluate 30 reviewer configurations. The benchmark reveals false robustness, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers differently. Moreover, the evaluated content-focused prompting protocol does not consistently improve robustness across backbones. Motivated by these findings, we introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a judgment based on an extracted, structured science core. This design combines manuscript-level assessment with a content-normalized view intended to reduce rhetorical sensitivity. In our primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile among the benchmarked reviewers while maintaining competitive human alignment. These results identify rhetorical robustness as a distinct evaluation target and demonstrate the potential of science-core review to improve it.