JudgeProfile: Understanding and Steering Subjectivity in LLM Judges

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inconsistency of preference in subjective comparisons by large language model (LLM) judges, proposing the JudgeProfile framework. This framework decouples the evaluation process into attribute perception and priority ranking, enabling standardized adaptation by revealing perceptual consensus and prioritization discrepancies. Furthermore, this work constructs the SubjectiveSet dataset and employs attribute weight estimation alongside reweighting techniques to guide judge decision-making, allowing adaptation to diverse preferences without fine-tuning. Experimental results demonstrate that the proposed method improves average consistency to 71.97%, significantly outperforming existing fine-tuning and prompt engineering approaches.
📝 Abstract
LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specific attributes like clarity, correctness, and detail) and prioritization (how much each attribute influences the final choice). We curate SubjectiveSet, a dataset of 50,013 response pairs from 17 public data sources, evaluated by 21 LLM judges across 87 attributes. We find a hidden consensus in perception: judges frequently agree on attribute judgments even when their overall choices diverge. Building on this separation, we first characterize each judge's prioritization using attribute weights estimated from its own overall choices. These weights differ across judges even when estimated from the same attribute judgments. We then learn new weights from reference labels to adapt their decisions to a target evaluation standard. Reweighting perceived attributes improves average held-out agreement with reference labels from 66.48% to 71.97%, outperforming fine-tuning and rubric prompting. Our findings show that understanding and steering the subjectivity of LLM judges requires attention not only to what they perceive, but also to how they prioritize it.
Problem

Research questions and friction points this paper is trying to address.

LLM judges
subjectivity
pairwise comparison
evaluation alignment
prioritization
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Judges
Subjectivity
JudgeProfile
Perception-Prioritization Decoupling
Attribute Reweighting
💼 Related Jobs
No related jobs found.
Q
Qi Cao
Meta
Kangning Liu
Kangning Liu
Adobe Inc
Machine LearningComputer Vision
Xuan Kan
Xuan Kan
Meta Platforms, Inc
Post TrainingLLMAI
S
Shunwen Tan
Meta
Y
Yang Pei
Meta
D
Dake Chen
Meta
Y
Yatai Ji
The University of Hong Kong
Z
Zixuan Ye
The Hong Kong University of Science and Technology
Y
Yuanpeng Tu
The University of Hong Kong
D
Daniel Li
Meta
J
Junbiao Tang
Meta
Pengtao Xie
Pengtao Xie
Associate Professor, UC San Diego; Adjunct Faculty, MBZUAI
Machine Learning
Zihao He
Zihao He
Meta
Natural Language ProcessingComputational Social ScienceAlignment