From Judgment Quality to Downstream Utility: Rethinking LLM-as-a-Judge for Open-Ended Tasks

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between evaluation quality and downstream utility in LLM-as-a-Judge paradigms for open-ended tasks, as well as their predominant confinement to the training phase. We systematically investigate how multi-dimensional judging protocols influence evaluation quality and innovatively extend the judge mechanism to the inference stage by integrating Best-of-N selection, guided revision, and beam search to optimize generation. Our findings reveal a notable discrepancy between intrinsic quality and downstream utility, demonstrating that protocol design substantially affects judging effectiveness. Furthermore, we validate that incorporating the judge into test-time computation effectively translates computational resources into tangible performance gains.
📝 Abstract
LLM-as-a-Judge is increasingly used to evaluate policy responses on open-ended tasks that lack ground-truth answers. Existing work often directly converts the resulting judgments into reward signals for policy training, paying limited attention to intrinsic judgment quality and largely restricting the use of Judges to training-time supervision. We systematically investigate judgment quality and downstream utility by examining both how judgments are elicited and how they are used. For judgment elicitation, we vary the Judge protocol along three dimensions: verdict granularity, critique usage, and evaluation batching. For judgment usage, beyond policy training, we extend Judge to test-time inference through Best-of-N selection, Judge-guided revision, and beam search. We find that, (i) Surprisingly, judgment quality and downstream utility do not always align. (ii) Judge protocol design substantially affects both intrinsic judgment quality and downstream utility. (iii) Judge guidance effectively converts test-time compute into performance gains, with benefits varying across inference strategies. Our results call for a multifaceted evaluation of LLM Judges on open-ended tasks, encompassing intrinsic judgment quality, and downstream utility.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-a-Judge
open-ended tasks
judgment quality
downstream utility
test-time inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-as-a-Judge
Test-time Inference
Judgment Quality
Downstream Utility
Judge Protocol
Z
Zheng Zhang
School of Information Science and Technology, ShanghaiTech University; State Key Laboratory of General Artificial Intelligence, BIGAI
L
Lufei Li
School of Information Science and Technology, ShanghaiTech University
X
Xinyue Tan
School of Information Science and Technology, ShanghaiTech University
Y
Yuanhao Zeng
School of Information Science and Technology, ShanghaiTech University; State Key Laboratory of General Artificial Intelligence, BIGAI
Z
Ziwei Shan
School of Information Science and Technology, ShanghaiTech University; State Key Laboratory of General Artificial Intelligence, BIGAI
Yexin Li
Yexin Li
State Key Laboratory of General Artificial Intelligence BIGAI
reinforcement learningmulti-agent systemmulti-armed banditsdata mining
Kan Ren
Kan Ren
Assistant Professor, ShanghaiTech University
Machine LearningData MiningLarge Language ModelFoundation Model