train with rubrics

Design, build, and analyze training and fine-tuning pipelines that use multi-criterion rubrics as structured supervision: convert rubric criteria into loss terms or reward signals, construct rubric-labeled example sets, and implement rubric-guided parameter or policy updates. Evaluate and iterate on methods that augment scalar rewards with multi-dimensional guidance, provide criterion-level feedback for policy updates, and improve learning and transfer for complex behaviors.

trainwithrubrics

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of structured mechanisms capable of dynamically evaluating and guiding behavior as large language models evolve toward open-ended autonomous agents. It proposes rubrics as a unified framework to translate complex quality judgments into structured, actionable specifications, systematically elucidating their progressive roles across evaluation, training, and internal agent behavior for the first time. By designing structured rubrics, decomposing assessments into multiple dimensions, generating dense feedback, and analyzing self-improvement behaviors, the study demonstrates the reliability of rubrics in ensuring generation quality, execution fidelity, adherence to theoretical constraints, and mitigation of safety threats. Furthermore, it establishes a cross-domain benchmarking framework that bridges human intent with machine behavior.

Autonomous AgentsEvaluation FrameworkHuman-AI Alignment

This work addresses the inefficiency of conventional reward-based reinforcement learning, which employs static weighting to aggregate human-defined criteria, conflating their prescribed importance with their actual utility during different training stages. To resolve this, the authors propose POW3R, a novel framework that introduces a policy-aware dynamic criterion weighting mechanism. POW3R adaptively amplifies reward signals from high-discriminability criteria based on policy output divergence, thereby decoupling the teaching signal from the final evaluation objective while preserving the latter. Built upon the GRPO algorithm and incorporating rollout-level contrastive analysis with multi-criterion reward modeling, POW3R significantly outperforms baselines across text and multimodal tasksโ€”winning 24 out of 30 comparisons across two datasets and three policies, achieving higher average rating rewards and strict completion rates, and reaching equivalent performance in 2.5โ€“4ร— fewer training steps.

policy-aware rewardsreinforcement learningreward aggregation

This work addresses the challenge of evaluating response quality in non-verifiable domainsโ€”such as creative writingโ€”where traditional reward models struggle to capture the multidimensional nature of output quality. To this end, the authors propose Rubric-ARM, a framework that integrates dynamic rubric generation with preference-based feedback, treating scoring rubrics as implicit actions. The approach jointly optimizes a rubric generator and a critic through reinforcement learning, employing an alternating training mechanism to effectively reduce gradient variance and enhance judgment accuracy. Experimental results demonstrate that Rubric-ARM achieves state-of-the-art performance across multiple benchmarks, significantly improving alignment between downstream policies and human preferences in both offline and online reinforcement learning settings.

LLM post-trainingnon-verifiable domainsresponse quality

This work addresses the challenge that open-ended generation tasks lack verifiable, fine-grained scoring rubrics, which limits the effectiveness of rule-based reinforcement learning. To overcome this, the authors propose an automated coarse-to-fine rubric generation framework that leverages principle-guided synthesis, multi-model aggregation, and a difficulty evolution mechanism to construct high-quality, highly discriminative scoring criteria. This framework enables the first large-scale, multi-domain, fine-grained, and scalable automatic evaluation system. Integrating the generated rubrics with Rejection Sampling Fine-Tuning (RuFT) and Rubric-guided Reinforcement Learning (RuRL), a Qwen3-14B model trained on RubricHub achieves a score of 69.3 on HealthBench, surpassing closed-source models such as GPT-5 and establishing a new state-of-the-art performance.

coarse criteriaopen-ended generationrubric-based evaluation

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing post-training methods for reasoning language models, which rely either on costly and noisy chain-of-thought annotations or on scalar rewards that lack fine-grained guidance. To overcome these challenges, the authors propose the first rubric-conditioned self-distillation framework that integrates task-specific, structured scoring rubrics into the training process, thereby providing token-level supervision signals for the reasoning trajectories generated by the student model and moving beyond the constraint of a single reference path. The approach employs a two-stage pipeline: first generating task-specific rubrics, then training the reasoning model via on-policy self-distillation augmented with fine-grained feedback derived from these rubrics. Evaluated across multiple scientific reasoning benchmarks, the method achieves consistent improvements, outperforming GRPO by 1.0 point and OPSD by 0.9 point on average, demonstrating a significant enhancement in reasoning capabilities.

chain-of-thoughtfine-grained feedbackreasoning language models

This work addresses the challenge of unreliable automatic evaluation signals in open-ended reasoning and long-form text generation, where conventional scoring rubrics struggle to capture knowledge-intensive dimensions, leading to distorted rewards. The authors propose DR-rubric, a two-stage framework that first employs multi-round agent-based search to uncover domain-specific facts, structural constraints, and failure modes, then distills these insights into atomic, independently verifiable constraints for GRPO policy optimization. Innovatively framing rubric construction as a dynamic research process, the approach replaces static templates with evidence-driven rule generation, enabling high-quality, self-bootstrapped scoring rules without reliance on state-of-the-art large language models. Evaluated across six benchmarks with only 1Kโ€“3K samples, the method significantly outperforms baselines: GPT-5-derived rules enhance coverage breadth, Gemini-based rules balance task performance, and iteratively refined self-bootstrapped rules achieve optimal overall results after three iterations.

long-form generationopen-ended reasoningpolicy optimization

Existing unsupervised scoring rule generation methods rely on a single evaluation persona, often overlooking critical dimensions of human preference and thereby introducing blind spots in assessment. This work proposes a Multi-Roles Scoring Rule Generation framework (MRRG), which introduces, for the first time, a training- and reference-free collaborative multi-role mechanism. By jointly generating and fusing interpretable scoring rules from complementary personas, MRRG enables verifiable pairwise preference validation and provides reward signals for reinforcement learning. Integrating multi-role prompting, rule fusion, and verifiable reward modeling, the method supports GRPO-style reinforcement learning. Empirical results demonstrate that MRRG significantly outperforms single-role baselines across multiple preference validation benchmarks, yielding more comprehensive and reliable reward signals that effectively enhance open-domain generation quality.

dimensional blind spotsLLM judgingpreference evaluation