🤖 AI Summary
This study challenges the prevailing assumption that iterative updates to large language models (LLMs) inherently enhance consistency in relevance judgment. Specifically, it investigates how backbone network evolution affects the stability of LLM-based evaluation. Employing the UMBRELA single-prompt and EXAM criteria-prompting methodologies, this work systematically assesses multiple generations of commercial and open-source models—including Gemini, GPT, Qwen, and Llama—under fixed prompting conditions. The findings reveal that improvements in aggregated performance do not necessarily translate to greater judgment stability. Notably, newer model versions fail to significantly improve relevance judgment quality and exhibit regression phenomena, wherein correct assessments made by older versions are lost in subsequent iterations. By exposing the latent risks associated with cross-version prompt transferability, this research provides a critical cautionary insight for LLM evaluation practices.
📝 Abstract
LLMs are evolving rapidly, with newer models offering stronger capabilities. This suggests that in LLM-based relevance judging, more capable models will achieve higher agreement with human judgements under the same prompt. We challenge this understanding by investigating the behavior of LLM-based relevance judges under backbone evolution. Keeping the prompts fixed, we evaluate a representative single-prompt (UMBRELA) and a rubric-based prompt (EXAM) across sequential model versions of commercial (Gemini, GPT) and open-weight (Qwen, Llama) models. Overall, we find no consistent evidence that newer versions lead to better relevance judges. Crucially, similar or improved aggregate performance does not imply judgment stability: correct judgements made by an earlier version of an LLM backbone are not necessarily preserved by later versions. We investigate the potential drivers of these regressions. Our findings caution against the assumption that judging prompts designed and validated for one backbone version will perform equivalently or better when the model is updated, even within the same family.