Score
Design expert rubrics is the competence of creating, synthesizing, and analyzing structured evaluation instruments that decompose prompts or tasks into atomic, intent-aware criteria with explicit scoring scales and calibration procedures, translating qualitative aspects of performance into unambiguous rubric items. It also includes producing question- or item-specific rubrics, automating rubric generation or synthesis where appropriate, and validating or calibrating rubrics (including with automated judges) to ensure context-sensitive coverage and reliable, high‑fidelity evaluation signals.
This work addresses the challenge of unreliable automatic evaluation signals in open-ended reasoning and long-form text generation, where conventional scoring rubrics struggle to capture knowledge-intensive dimensions, leading to distorted rewards. The authors propose DR-rubric, a two-stage framework that first employs multi-round agent-based search to uncover domain-specific facts, structural constraints, and failure modes, then distills these insights into atomic, independently verifiable constraints for GRPO policy optimization. Innovatively framing rubric construction as a dynamic research process, the approach replaces static templates with evidence-driven rule generation, enabling high-quality, self-bootstrapped scoring rules without reliance on state-of-the-art large language models. Evaluated across six benchmarks with only 1K–3K samples, the method significantly outperforms baselines: GPT-5-derived rules enhance coverage breadth, Gemini-based rules balance task performance, and iteratively refined self-bootstrapped rules achieve optimal overall results after three iterations.
This study addresses the lack of systematic statistical analysis on how modifications to scoring rubrics affect agreement between human raters and large language model (LLM)-based automated scoring systems. It presents the first quantitative investigation into the impact of various rubric design choices—including the incorporation of contextual exemplars, adjustments in complexity, and mitigation of position bias—on human–LLM scoring consistency. Drawing on data from automated essay scoring and instruction-following evaluation tasks, the authors employ statistical methods to compare outcomes under holistic versus analytic rubric configurations. Their findings reveal that integrating representative examples and reducing position bias significantly enhance inter-rater consistency, whereas highly complex rubrics and conservative aggregation strategies tend to diminish it. These results offer empirical guidance for optimizing rubric design in automated scoring systems.
This work addresses the lack of structured mechanisms capable of dynamically evaluating and guiding behavior as large language models evolve toward open-ended autonomous agents. It proposes rubrics as a unified framework to translate complex quality judgments into structured, actionable specifications, systematically elucidating their progressive roles across evaluation, training, and internal agent behavior for the first time. By designing structured rubrics, decomposing assessments into multiple dimensions, generating dense feedback, and analyzing self-improvement behaviors, the study demonstrates the reliability of rubrics in ensuring generation quality, execution fidelity, adherence to theoretical constraints, and mitigation of safety threats. Furthermore, it establishes a cross-domain benchmarking framework that bridges human intent with machine behavior.
This work addresses the limitations of existing large language model (LLM) evaluation methods, which struggle to effectively assess complex, context-dependent instruction following and agent behaviors, often focusing only on superficial constraints. To overcome this, we propose a novel evaluation paradigm grounded in expert-designed rubrics, featuring atomic, intent-aware scoring criteria calibrated via LLM-based judges to enable precise assessment and efficient training on complex tasks. We introduce five principles for high-quality rubric design and, for the first time, leverage expert rubrics simultaneously as both evaluation instruments and reinforcement learning signals. Experiments demonstrate that models trained on our ComplexConstraints dataset exhibit substantial improvements—15.5% and 12.2% gains in instruction-following performance for 4B and 235B parameter models, respectively—and show strong generalization to unseen enterprise-level tasks, with notable improvements on BFCL (+4.5%), Tau2-Bench (+7.4%), and Tool-Decathlon (+6.8%).
This study addresses the high cost of manually constructing reproducibility scoring rubrics, which hinders the scalability of benchmarks like PaperBench. It presents the first systematic evaluation of the reliability of large language models (LLMs) in automatically generating such rubrics, reformulated as checklists, using two backbone LLMs across four generation settings. Through internal and external meta-evaluations based on semantic similarity and alignment with human benchmarks, the authors find that enhanced generation strategies significantly improve downstream evaluation alignment, with the best configuration approaching human-level performance. Nevertheless, limitations persist in fine-grained coverage, scoring bias, and domain adaptability, and improvements in intrinsic semantic quality remain modest.
This work addresses the limitations of existing rubric-based evaluation methods, which flatten scoring criteria into simplistic prompts and fail to capture the compositional logic among evaluation dimensions. To overcome this, the authors propose formalizing rubrics as task-agnostic, typed directed evaluation graphs that enable structured scoring through criterion nodes, transformation/reduction/gating operators, and task-specific readout functions. The approach introduces, for the first time, a static type system and port-based connectivity mechanism, unifying support for both pointwise and pairwise evaluation tasks while enabling end-to-end reasoning with large language models. Evaluated on GPT-OSS-120B, the method improves Exact Score Agreement by 0.62–6.75 percentage points in pointwise assessment and achieves state-of-the-art end-to-end accuracy on two benchmarks for pairwise preference judgment.
This work addresses the lack of systematic evaluation of rubric reliability when using large language models as judges (LLM-as-a-Judge) in complex, long-form agent tasks. The authors propose RuVerBench, the first benchmark specifically designed for rubric validation in agent-centric scenarios, spanning deep research and agent programming domains. Through a large-scale human-annotated dataset, multi-model comparisons, and prompt engineering analyses, they systematically investigate the impact of prompt design, batch processing, and majority voting strategies. Experimental results show that while state-of-the-art models generally perform well, substantial noise remains; weaker models exhibit greater sensitivity to prompt variations, batch processing entails trade-offs between accuracy and efficiency, and majority voting improves reliability but with diminishing returns.