Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges

📅 2026-08-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing rubric-based evaluation methods, which flatten scoring criteria into simplistic prompts and fail to capture the compositional logic among evaluation dimensions. To overcome this, the authors propose formalizing rubrics as task-agnostic, typed directed evaluation graphs that enable structured scoring through criterion nodes, transformation/reduction/gating operators, and task-specific readout functions. The approach introduces, for the first time, a static type system and port-based connectivity mechanism, unifying support for both pointwise and pairwise evaluation tasks while enabling end-to-end reasoning with large language models. Evaluated on GPT-OSS-120B, the method improves Exact Score Agreement by 0.62–6.75 percentage points in pointwise assessment and achieves state-of-the-art end-to-end accuracy on two benchmarks for pairwise preference judgment.
📝 Abstract
Rubric-based evaluators commonly treat rubrics as prompt context or flat criteria: they specify what to judge but leave criterion composition implicit, even when natural-language rules state it. We introduce Graph-Structured Rubrics (GSR), which compiles a rubric into a response-independent typed evaluation graph before observing responses. Criterion nodes elicit judgments; transformation, reduction, and gating operators compose them through named ports; and a task-specific output mapping, termed Readout, converts the unique sink into a score or preference. Compilation rejects malformed or type-incompatible graphs. Pointwise evaluation judges rubric dimensions separately before graph aggregation; pairwise evaluation reuses the graph with one judgment for each candidate under every criterion. Under GPT-OSS-120B, GSR improves exact score agreement by 0.62--6.75 percentage points over Prometheus-style scoring on four pointwise datasets and achieves the numerically highest end-to-end pairwise accuracy on two preference benchmarks under native tie and abstention policies.
Problem

Research questions and friction points this paper is trying to address.

rubric-based evaluation
criterion composition
evaluation graphs
LLM judges
structured scoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph-Structured Rubrics
Typed Evaluation Graph
LLM Judges
Rubric Compilation
Readout Mapping
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xi Chen
Ant Group
J
Jie Mu
Ant Group
M
Mo Xuan
Ant Group
Q
Qun Shao
Ant Group