🤖 AI Summary
This study addresses the inconsistency of preference in subjective comparisons by large language model (LLM) judges, proposing the JudgeProfile framework. This framework decouples the evaluation process into attribute perception and priority ranking, enabling standardized adaptation by revealing perceptual consensus and prioritization discrepancies. Furthermore, this work constructs the SubjectiveSet dataset and employs attribute weight estimation alongside reweighting techniques to guide judge decision-making, allowing adaptation to diverse preferences without fine-tuning. Experimental results demonstrate that the proposed method improves average consistency to 71.97%, significantly outperforming existing fine-tuning and prompt engineering approaches.
📝 Abstract
LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specific attributes like clarity, correctness, and detail) and prioritization (how much each attribute influences the final choice). We curate SubjectiveSet, a dataset of 50,013 response pairs from 17 public data sources, evaluated by 21 LLM judges across 87 attributes. We find a hidden consensus in perception: judges frequently agree on attribute judgments even when their overall choices diverge. Building on this separation, we first characterize each judge's prioritization using attribute weights estimated from its own overall choices. These weights differ across judges even when estimated from the same attribute judgments. We then learn new weights from reference labels to adapt their decisions to a target evaluation standard. Reweighting perceived attributes improves average held-out agreement with reference labels from 66.48% to 71.97%, outperforming fine-tuning and rubric prompting. Our findings show that understanding and steering the subjectivity of LLM judges requires attention not only to what they perceive, but also to how they prioritize it.