CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Reliable evaluation of open-domain large language models demands fine-grained scoring rubrics, yet expert-authored rubrics are costly to construct, and existing automated methods struggle to identify measurable and informative scoring items. This work proposes a task-adaptive framework for building scoring rubric libraries by integrating typed scoring, Bayesian measurability filtering, and item response theory (IRT)-driven submodular optimization. The approach estimates the measurability of scoring items via Beta-Bernoulli posterior inference and performs uncertainty-aware selection by optimizing submodular information coverage. Evaluated on JudgmentBench, the method improves agreement with human gold-standard judgments (Cohen’s κ) from 0.604 to 0.743. On FinResearchBench, it achieves the same correlation target using only 49 rubric items compared to the original 131, significantly outperforming baselines in cross-task ranking fidelity.
📝 Abstract
Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge unanimity and binary variance filters, which cannot distinguish measurable rubrics from informative ones. We introduce CalibratedRubric, a task-adaptive framework that combines type-specific scoring, Bayesian rubric-measurability filtering, and item response theory (IRT)-based bank assembly. CalibratedRubric estimates each rubric's measurability with a Beta--Bernoulli agreement posterior and uses a submodular information-coverage objective to construct compact rubric banks over the observed capability range. Across financial, healthcare, general, and legal benchmarks, measurability filtering improves human-gold agreement on JudgmentBench from $κ=0.604$ to $0.743$. IRT-based greedy selection improves cross-fitted rank fidelity over random selection across all six evaluated response blocks and requires only 49 rather than 131 rubrics to reach the target correlation on FinResearchBench decision-support tasks. Task-label perturbations further reduce system separation, confirming the practical relevance of task-adaptive scoring. These results support CalibratedRubric as an efficient, uncertainty-aware approach to open-ended LLM evaluation, with calibration gains depending on sufficient judge redundancy.
Problem

Research questions and friction points this paper is trying to address.

open-ended LLM evaluation
rubric calibration
measurability filtering
task-adaptive scoring
evaluation reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

task-adaptive rubrics
Bayesian measurability filtering
item response theory
submodular optimization
LLM evaluation
🔎 Similar Papers
No similar papers found.
Mengting Chen
Mengting Chen
Alibaba Group
Generative ModelingComputer Vision
Y
Yanshu Sun
FinStep
W
Wanting Liang
FinStep
B
Beidi Luan
StepFun
R
Rui Sun
StepFun
Dezhi Chen
Dezhi Chen
Beijing University of Posts and Telecommunications
UAV,Game theory,networks,Mean field
J
Jing Li
StepFun
Z
Zuo Bai
StepFun, FinStep