Evaluating Rubric Generation with Interventional Transfer

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of scalably validating the quality of evaluation criteria generated by large language models (LLMs). To this end, we propose an intervention transfer method that dynamically measures the utility similarity between LLM-generated and expert-defined criteria by perturbing model responses and observing the resulting co-variation in scores. This approach reveals asymmetric transfer phenomena among criteria, thereby overcoming the limitations of traditional static comparisons. Through multi-model comparative experiments on the HealthBench benchmark, we empirically demonstrate the asymmetry inherent in criteria generated by different models. These findings offer a novel paradigm for optimizing performance monitoring strategies in LLM-based evaluation pipelines.
📝 Abstract
Instance-specific rubrics are common in AI benchmarks where reliable evaluation requires specific expert knowledge. This approach is difficult to scale, prompting research into the generation of rubrics with large language models (LLMs). However, even when expert rubrics are available as references, it is unclear how to productively evaluate the quality of generated rubrics at scale. In this paper, we introduce a method for the evaluation of rubric generation, which we call Interventional Transfer (IT), based on the idea that two rubrics are similar if they move together when a response is perturbed to pass/fail one of them. In contrast to existing approaches for evaluation of rubric generation, we argue that different forms of interventional transfer can be used to evaluate the utility of generated rubrics for different tasks. For instance, we apply this approach in a case study on HealthBench, where we demonstrate an asymmetry in rubrics generated by Qwen3.8-27B, Deepseek-V4-Flash, and Opus-5, used to evaluate responses from GPT-5.6-Terra. Perturbations that degrade responses according to the generated/expert rubric transfer into lower scores on the corresponding expert/generated rubric, but perturbations that improve on one rubric do not reliably transfer into higher scores on the other. We argue that this finding has implications for the usage of LLM-generated rubrics for performance monitoring and hill-climbing. We contrast our approach with existing approaches for rubric evaluation, which do not surface the same asymmetry that we observe.
Problem

Research questions and friction points this paper is trying to address.

Rubric Generation
Evaluation Quality
Large Language Models
AI Benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interventional Transfer
Rubric Generation
Large Language Models
Evaluation Methodology
Asymmetry
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
E
Erik Skalnes
Johns Hopkins Computer Science
L
Layne C. Price
Amazon
Raviteja Anantha
Raviteja Anantha
Amazon
Deep LearningGenerative AIAI for Drug DiscoveryGenomicsPersonalized Medicine
M
Michael Oberst
Johns Hopkins Computer Science