🤖 AI Summary
This study addresses the illusion of reliability and insufficient vulnerability coverage in LLM review evaluation caused by static templates by constructing a three-tiered assessment framework to reveal its hierarchical fragility. We innovatively propose SCOPE-Fuzzer, a dynamic probing mechanism that introduces strategy-aware fuzzing into this domain for the first time. By integrating feedback-driven strategy selection with adaptive content mutation, this approach overcomes the limitations of traditional static evaluation. Experimental results demonstrate that the proposed method consistently uncovers latent review vulnerabilities overlooked by static assessments and baseline techniques, significantly enhancing evaluation effectiveness and robustness.
📝 Abstract
The rapid growth of submissions and reviewing workload has accelerated the use of large language models (LLMs) in peer review. Prior studies suggest that LLM-based reviewers can penalize content perturbations, such as overclaiming, indicating a certain degree of reliability. Yet these conclusions are largely based on a narrow set of perturbation strategies instantiated with static templates, providing limited evidence of actual reliability. In this paper, we construct a three-level evaluation framework covering perturbations to surface presentation, argumentative logic, and value judgment. Experiments on representative LLM-based reviewers reveal two limitations of static evaluation: stratified vulnerability, where perturbation effects depend on whether the paper's original review score is high or low, and perturbation undercoverage, where a single template misses vulnerabilities exposed by diverse realizations. To address these limitations, we propose SCOPE-Fuzzer, a strategy-aware fuzzer that combines feedback-driven strategy selection with adaptive mutation of paper content. By iteratively probing reviewers with dynamic perturbations, SCOPE-Fuzzer consistently uncovers vulnerabilities overlooked by static evaluation and other baselines.