LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

📅 2026-07-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that existing evaluation metrics for literature review agents fail to reflect expert judgment and research utility. To this end, we construct a competitive evaluation platform featuring the first battle protocol and five-dimensional criteria tailored for literature reviews, collecting multidimensional preference data through anonymous pairwise comparisons by domain experts. Building upon this, we propose LitJudge, an expert-calibrated evaluator that integrates human-AI collaborative review with LLM-as-a-judge alignment techniques to mitigate judge bias. Experimental results demonstrate that even top-performing systems achieve only a 23% win rate, while LitJudge improves evaluation consistency to a Spearman correlation of 0.78, closely approximating human expert performance.
📝 Abstract
Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM-as-a-judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis-heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert-calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter-expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.
Problem

Research questions and friction points this paper is trying to address.

literature review evaluation
automated literature review
LLM-as-a-judge
expert judgment
Innovation

Methods, ideas, or system contributions that make the work stand out.

LitReview Arena
Battle-style evaluation
Literature review agents
LitJudge
Expert-calibrated evaluator
R
Ruotong Zhao
Tsinghua University, Beijing, China
Zhiyu Chen
Zhiyu Chen
Amazon
Conversational AILarge Language ModelsInformation RetrievalNatural language Processing
X
Xurui Liu
Tsinghua University, Beijing, China
H
Haidong Xue
Zhongguancun Institute of Artificial Intelligence, Beijing, China
D
Dong Liang
Zhongguancun Institute of Artificial Intelligence, Beijing, China
J
Jigao Fu
Zhongguancun Institute of Artificial Intelligence, Beijing, China
Y
Yanbiao Wu
Zhongguancun Institute of Artificial Intelligence, Beijing, China
Y
Yuanyi Zhen
Zhongguancun Academy, Beijing, China
Fengli Xu
Fengli Xu
Tsinghua University
LLM AgentData ScienceSocial ComputingScience of ScienceUrban Science
Y
Yong Li
Tsinghua University, Beijing, China