Breaking the Illusion of Review Reliability under Static Evaluation: SCOPE Fuzzing for LLM-based Scientific Reviewers

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the illusion of reliability and insufficient vulnerability coverage in LLM review evaluation caused by static templates by constructing a three-tiered assessment framework to reveal its hierarchical fragility. We innovatively propose SCOPE-Fuzzer, a dynamic probing mechanism that introduces strategy-aware fuzzing into this domain for the first time. By integrating feedback-driven strategy selection with adaptive content mutation, this approach overcomes the limitations of traditional static evaluation. Experimental results demonstrate that the proposed method consistently uncovers latent review vulnerabilities overlooked by static assessments and baseline techniques, significantly enhancing evaluation effectiveness and robustness.
📝 Abstract
The rapid growth of submissions and reviewing workload has accelerated the use of large language models (LLMs) in peer review. Prior studies suggest that LLM-based reviewers can penalize content perturbations, such as overclaiming, indicating a certain degree of reliability. Yet these conclusions are largely based on a narrow set of perturbation strategies instantiated with static templates, providing limited evidence of actual reliability. In this paper, we construct a three-level evaluation framework covering perturbations to surface presentation, argumentative logic, and value judgment. Experiments on representative LLM-based reviewers reveal two limitations of static evaluation: stratified vulnerability, where perturbation effects depend on whether the paper's original review score is high or low, and perturbation undercoverage, where a single template misses vulnerabilities exposed by diverse realizations. To address these limitations, we propose SCOPE-Fuzzer, a strategy-aware fuzzer that combines feedback-driven strategy selection with adaptive mutation of paper content. By iteratively probing reviewers with dynamic perturbations, SCOPE-Fuzzer consistently uncovers vulnerabilities overlooked by static evaluation and other baselines.
Problem

Research questions and friction points this paper is trying to address.

LLM-based scientific reviewers
static evaluation
peer review reliability
vulnerability detection
content perturbation
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-based Reviewers
Fuzzing
Dynamic Perturbation
Evaluation Framework
Adaptive Mutation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Zhuo Chen
Zhuo Chen
Wuhan University
LLM securityRAG reliabilityInformation retireval
H
Hao Zeng
Wuhan University
Jiawei Liu
Jiawei Liu
Wuhan University
Information RetrievalContent SecurityDocument Intelligence
G
Guoxiu He
East China Normal University
L
Le Cai
Wuhan University
L
Liu Haotan
Wuhan University
L
Li Wenbo
Wuhan University
Y
Yong Huang
Wuhan University
Wei Lu
Wei Lu
Professor of Information Management, Wuhan University
Information RetrievalGraph theorySocial networksWeb metrics