design expert rubrics

Design expert rubrics is the competence of creating, synthesizing, and analyzing structured evaluation instruments that decompose prompts or tasks into atomic, intent-aware criteria with explicit scoring scales and calibration procedures, translating qualitative aspects of performance into unambiguous rubric items. It also includes producing question- or item-specific rubrics, automating rubric generation or synthesis where appropriate, and validating or calibrating rubrics (including with automated judges) to ensure context-sensitive coverage and reliable, high‑fidelity evaluation signals.

designexpertrubrics

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.94
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$210K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of unreliable automatic evaluation signals in open-ended reasoning and long-form text generation, where conventional scoring rubrics struggle to capture knowledge-intensive dimensions, leading to distorted rewards. The authors propose DR-rubric, a two-stage framework that first employs multi-round agent-based search to uncover domain-specific facts, structural constraints, and failure modes, then distills these insights into atomic, independently verifiable constraints for GRPO policy optimization. Innovatively framing rubric construction as a dynamic research process, the approach replaces static templates with evidence-driven rule generation, enabling high-quality, self-bootstrapped scoring rules without reliance on state-of-the-art large language models. Evaluated across six benchmarks with only 1K–3K samples, the method significantly outperforms baselines: GPT-5-derived rules enhance coverage breadth, Gemini-based rules balance task performance, and iteratively refined self-bootstrapped rules achieve optimal overall results after three iterations.

long-form generationopen-ended reasoningpolicy optimization

This study addresses the lack of systematic statistical analysis on how modifications to scoring rubrics affect agreement between human raters and large language model (LLM)-based automated scoring systems. It presents the first quantitative investigation into the impact of various rubric design choices—including the incorporation of contextual exemplars, adjustments in complexity, and mitigation of position bias—on human–LLM scoring consistency. Drawing on data from automated essay scoring and instruction-following evaluation tasks, the authors employ statistical methods to compare outcomes under holistic versus analytic rubric configurations. Their findings reveal that integrating representative examples and reducing position bias significantly enhance inter-rater consistency, whereas highly complex rubrics and conservative aggregation strategies tend to diminish it. These results offer empirical guidance for optimizing rubric design in automated scoring systems.

analytic judgmentholistic judgmenthuman-autorater agreement

This work addresses the lack of structured mechanisms capable of dynamically evaluating and guiding behavior as large language models evolve toward open-ended autonomous agents. It proposes rubrics as a unified framework to translate complex quality judgments into structured, actionable specifications, systematically elucidating their progressive roles across evaluation, training, and internal agent behavior for the first time. By designing structured rubrics, decomposing assessments into multiple dimensions, generating dense feedback, and analyzing self-improvement behaviors, the study demonstrates the reliability of rubrics in ensuring generation quality, execution fidelity, adherence to theoretical constraints, and mitigation of safety threats. Furthermore, it establishes a cross-domain benchmarking framework that bridges human intent with machine behavior.

Autonomous AgentsEvaluation FrameworkHuman-AI Alignment

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing large language model (LLM) evaluation methods, which struggle to effectively assess complex, context-dependent instruction following and agent behaviors, often focusing only on superficial constraints. To overcome this, we propose a novel evaluation paradigm grounded in expert-designed rubrics, featuring atomic, intent-aware scoring criteria calibrated via LLM-based judges to enable precise assessment and efficient training on complex tasks. We introduce five principles for high-quality rubric design and, for the first time, leverage expert rubrics simultaneously as both evaluation instruments and reinforcement learning signals. Experiments demonstrate that models trained on our ComplexConstraints dataset exhibit substantial improvements—15.5% and 12.2% gains in instruction-following performance for 4B and 235B parameter models, respectively—and show strong generalization to unseen enterprise-level tasks, with notable improvements on BFCL (+4.5%), Tau2-Bench (+7.4%), and Tool-Decathlon (+6.8%).

agentic taskscomplex constraintsinstruction following

This study addresses the high cost of manually constructing reproducibility scoring rubrics, which hinders the scalability of benchmarks like PaperBench. It presents the first systematic evaluation of the reliability of large language models (LLMs) in automatically generating such rubrics, reformulated as checklists, using two backbone LLMs across four generation settings. Through internal and external meta-evaluations based on semantic similarity and alignment with human benchmarks, the authors find that enhanced generation strategies significantly improve downstream evaluation alignment, with the best configuration approaching human-level performance. Nevertheless, limitations persist in fine-grained coverage, scoring bias, and domain adaptability, and improvements in intrinsic semantic quality remain modest.

benchmark scalabilityLLM evaluationopen-ended output assessment

This work addresses the limitations of existing rubric-based evaluation methods, which flatten scoring criteria into simplistic prompts and fail to capture the compositional logic among evaluation dimensions. To overcome this, the authors propose formalizing rubrics as task-agnostic, typed directed evaluation graphs that enable structured scoring through criterion nodes, transformation/reduction/gating operators, and task-specific readout functions. The approach introduces, for the first time, a static type system and port-based connectivity mechanism, unifying support for both pointwise and pairwise evaluation tasks while enabling end-to-end reasoning with large language models. Evaluated on GPT-OSS-120B, the method improves Exact Score Agreement by 0.62–6.75 percentage points in pointwise assessment and achieves state-of-the-art end-to-end accuracy on two benchmarks for pairwise preference judgment.

criterion compositionevaluation graphsLLM judges

This work addresses the lack of systematic evaluation of rubric reliability when using large language models as judges (LLM-as-a-Judge) in complex, long-form agent tasks. The authors propose RuVerBench, the first benchmark specifically designed for rubric validation in agent-centric scenarios, spanning deep research and agent programming domains. Through a large-scale human-annotated dataset, multi-model comparisons, and prompt engineering analyses, they systematically investigate the impact of prompt design, batch processing, and majority voting strategies. Experimental results show that while state-of-the-art models generally perform well, substantial noise remains; weaker models exhibit greater sensitivity to prompt variations, batch processing entails trade-offs between accuracy and efficiency, and majority voting improves reliability but with diminishing returns.

agentic scenariosLLM-as-a-Judgemodel evaluation

Hot Scholars

RS

Renee Shelby

Staff Research Scientist, Google Research
AI social impactstechnologyinequalityfeminism
YZ

Yuanxing Zhang

Kuaishou Technology
Recommender SystemLarge Language ModelVideo Understanding
MM

Mantas Mazeika

Center for AI Safety
ML SafetyAI SafetyMachine EthicsML Reliability
YH

Yunzhong He

University of California, Los Angeles
machine learningnatural language processinginformation retrievalrobot learning