Rubric-Calibrated Preferences: Cross-Query Calibration of LLM Judgments via Item Response Theory

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of sparse and noisy human labels, insufficient nDCG discriminability, and cross-query incomparability of LLM-based scores in reranking evaluation. We propose an assessment framework grounded in listwise tournaments and rigorous criteria, introducing a shared rubric to unify the relative and absolute judgment scales of LLMs for the first time. By leveraging Item Response Theory (IRT), we calibrate relative judgments into dense, cross-query comparable relevance probabilities, and introduce the RCP-nDCG metric to replace traditional discrete labels. This approach significantly enhances fine-grained model discrimination, achieving a correlation coefficient of 0.795 with human ratings and an AUC of 0.910. Furthermore, 72.4% of pairwise comparisons align with expert assessments, substantially reducing evaluation ties.
📝 Abstract
Rerankers decide which documents users and LLMs see, yet their standard metric, nDCG, relies on human relevance labels that are costly, sparse, noisy, and discretely graded. As rerankers approach each other in quality, nDCG on these labels therefore increasingly fails to separate them. LLM judges could supply dense labels. Relative judgments within one query tell even close candidates apart, yet their scores share no scale across queries. Absolute grades share one scale but are too coarse to distinguish documents of similar relevance. We propose Rubric-Calibrated Preferences (RCP), which combine both kinds of judgment. A listwise Bradley-Terry tournament orders each query's documents, and a rubric of yes/no criteria of increasing stringency provides an absolute standard. Item Response Theory (IRT), which scores test-takers based on their answers to common questions, then uses the shared criteria to put all queries'tournament scores on one scale. RCP's retrieval metric, RCP-nDCG, replaces nDCG's discrete labels with the resulting calibrated relevance probabilities. Against blind grades from 46 external annotators, calibration raises the correlation between a query's mean score and its mean human grade from 0.538 to 0.795. The probabilities rank a useful document above a non-useful one with probability 0.910 (AUC, chance 0.5), versus 0.651 for the benchmark labels. When the annotators'grades prefer one of two rerankers and exactly one metric agrees, that metric is RCP-nDCG in 72.4% of 185 comparisons (chance about 53%). On TREC-DL, RCP-nDCG sides with NIST assessors'grades on every reranker pair that these grades separate significantly. RCP-nDCG also resolves many of nDCG's ties and separates 1.9 times as many reranker pairs on NanoBEIR. Rubric calibration thus turns relative LLM judgments into dense relevance labels that are comparable across queries and agree with human judgment.
Problem

Research questions and friction points this paper is trying to address.

reranker evaluation
LLM judges
nDCG
relevance labels
cross-query calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Item Response Theory
Rubric-Calibrated Preferences
LLM Judges
Reranker Evaluation
Bradley-Terry Model