MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of "hollow credit" in rubric-based reinforcement learning, where judges erroneously reward responses lacking critical information. To mitigate this, we propose MetaRubric, a framework integrated with the GRPO algorithm that introduces counterfactual data construction and dynamic rubric revision. By alternating between evidence-aware policy optimization and response-guided rubric adaptation, the method ensures rewards are allocated only when responses provide sufficient evidence, while adaptively adjusting weights to correct policy bias. Experimental results demonstrate that MetaRubric outperforms static baselines by 6.00 to 20.40 percentage points on PubMedQA and achieves substantial improvements on HealthBench-Hard and multimodal medical benchmarks.
📝 Abstract
Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the information or action it requires is absent from the response, a failure mode we term Vacuous Credit. Such awards persist after the required information is removed and can reverse the sign of a response's GRPO advantage. To address this problem, we introduce MetaRubric, which alternates evidence-aware policy optimization with response-guided rubric adaptation. We construct counterfactual counterparts by changing one task-relevant fact in each prompt. During policy optimization, credit is assigned only when the response contains sufficient evidence to satisfy the required rubric criterion. After each policy-optimization stage, current policy responses guide revisions to original and counterfactual criteria while preserving the meaning of the original prompt's initial rubric as interpreted under each prompt's facts. We also adapt criterion weights at stage boundaries to better address observed policy errors. Across multiple backbones, MetaRubric improves PubMedQA accuracy by 6.00--20.40 percentage points over static-judge GRPO, with further gains on HealthBench-Hard and two multimodal medical benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Rubric-based Reinforcement Learning
Vacuous Credit
Reward Modeling
GRPO
Criterion Scoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rubric-based reinforcement learning
Vacuous Credit
Counterfactual construction
Evidence-aware policy optimization
Rubric adaptation