Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that relying solely on overall score alignment when using large language models (LLMs) as automated judges fails to reveal discrepancies between LLM and human evaluations on difficult cases. To investigate this, we introduce the psychometric Many-Facet Rasch Model and propose the concept of "residual difficulty" to decompose scoring structures. Using the SummEval dataset with 17 open-source LLM judges, we systematically compare residual difficulty distributions across evaluation dimensions between humans and machines. Our findings demonstrate that overall score agreement obscures dimension-dependent misalignments in evaluation difficulty; for instance, LLMs struggle more with assessing consistency, whereas humans find coherence harder to evaluate. This work provides both a theoretical foundation and methodological support for developing more precise human-LLM collaborative evaluation frameworks.
📝 Abstract
Large language models (LLMs) are widely used as automatic judges, with validity typically assessed via alignment with human scores. However, aggregate agreement fails to reveal whether humans and LLMs find the same evaluation cases difficult. In this paper, we study this problem in summarization evaluation from a psychometric perspective. We fit Many-Facet Rasch Models separately to human and LLM ratings to decompose scores into latent summary quality, rater severity, dimension severity, and rating-scale thresholds. Building on this decomposition, we define residual hardness as a model-adjusted measure of judging difficulty and compare whether human and LLM judges share the same hardness structure. Across 17 open-weight LLM judges on SummEval, we find that moderate alignment in latent summary quality does not imply alignment in residual hardness. Human and LLM judges differ in which summary--dimension units remain difficult, and this mismatch is strongly dimension-dependent. Consistency shows a pronounced LLM-hard shift, whereas coherence shows a human-hard shift. We further show that human-easy but LLM-hard cases are partially predictable from observable source--summary properties. These findings suggest that aggregate human alignment reflects only part of LLM-as-a-judge reliability, while psychometric residual diagnostics support more informative judge evaluation and more targeted human--LLM collaboration.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-a-Judge
Psychometric Analysis
Residual Judging Difficulty
Summarization Evaluation
Score Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Psychometric Analysis
Many-Facet Rasch Models
Residual Hardness
LLM-as-a-Judge
Summarization Evaluation
💼 Related Jobs
No related jobs found.
L
Longwei Cong
DIPF | Leibniz Institute for Research and Information in Education
S
Sonja Hahn
DIPF | Leibniz Institute for Research and Information in Education
S
Sebastian Gombert
DIPF | Leibniz Institute for Research and Information in Education
L
Leon Camus
DIPF | Leibniz Institute for Research and Information in Education
Fabian Zehner
Fabian Zehner
DIPF | Leibniz Institute for Research and Information in Education; Centre for International Student Assessment (ZIB)
Hendrik Drachsler
Hendrik Drachsler
Professor for Computer Science, DIPF | Leibniz Institute & Goethe University, Frankfurt
Learning AnalyticsAI in EducationAssessment and FeedbackLearning DesignMedical Education
U
Ulf Kroehne
DIPF | Leibniz Institute for Research and Information in Education; Chemnitz University of Technology