🤖 AI Summary
This study addresses the limitation that relying solely on overall score alignment when using large language models (LLMs) as automated judges fails to reveal discrepancies between LLM and human evaluations on difficult cases. To investigate this, we introduce the psychometric Many-Facet Rasch Model and propose the concept of "residual difficulty" to decompose scoring structures. Using the SummEval dataset with 17 open-source LLM judges, we systematically compare residual difficulty distributions across evaluation dimensions between humans and machines. Our findings demonstrate that overall score agreement obscures dimension-dependent misalignments in evaluation difficulty; for instance, LLMs struggle more with assessing consistency, whereas humans find coherence harder to evaluate. This work provides both a theoretical foundation and methodological support for developing more precise human-LLM collaborative evaluation frameworks.
📝 Abstract
Large language models (LLMs) are widely used as automatic judges, with validity typically assessed via alignment with human scores. However, aggregate agreement fails to reveal whether humans and LLMs find the same evaluation cases difficult. In this paper, we study this problem in summarization evaluation from a psychometric perspective. We fit Many-Facet Rasch Models separately to human and LLM ratings to decompose scores into latent summary quality, rater severity, dimension severity, and rating-scale thresholds. Building on this decomposition, we define residual hardness as a model-adjusted measure of judging difficulty and compare whether human and LLM judges share the same hardness structure. Across 17 open-weight LLM judges on SummEval, we find that moderate alignment in latent summary quality does not imply alignment in residual hardness. Human and LLM judges differ in which summary--dimension units remain difficult, and this mismatch is strongly dimension-dependent. Consistency shows a pronounced LLM-hard shift, whereas coherence shows a human-hard shift. We further show that human-easy but LLM-hard cases are partially predictable from observable source--summary properties. These findings suggest that aggregate human alignment reflects only part of LLM-as-a-judge reliability, while psychometric residual diagnostics support more informative judge evaluation and more targeted human--LLM collaboration.