OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of sparse reward signals and imprecise credit assignment in rubric-based reinforcement learning for open-ended tasks. To overcome these bottlenecks, we propose a two-stage training framework. First, we introduce Rubric-Privileged Online Policy Distillation (RP-OPD), which pioneers the use of evaluation rubrics as privileged teacher context to provide dense, token-level supervision. Subsequently, the model undergoes reinforcement learning by directly optimizing rubric-based rewards. This mechanism effectively mitigates reward hacking during the subsequent RL phase. Empirical evaluations demonstrate that our approach significantly outperforms SFT+RL baselines across multiple benchmarks, achieving state-of-the-art performance and enhancing the model’s genuine adherence to specified evaluation rubrics.
πŸ“ Abstract
Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this issue by scoring open-ended responses against explicit criteria. However, because the reward is assigned after the complete response, the training signal does not directly identify which individual decisions contributed to the final score. We propose a two-stage training framework that uses rubrics first as privileged teacher context for dense token-level supervision, then as rewards for further RL. In the first stage, rubric-privileged on-policy distillation (RP-OPD), a student without access to the rubric matches a rubric-aware teacher's next-token distributions at student-generated prefixes. In the second stage, RL directly optimizes the rubric reward and improves beyond the observed distillation plateau. We evaluate the framework on health and science tasks using open-weight models. Across HealthBench, ResearchQA, and RubricHub Science, we compare post-training methods and vary the amount of SFT or RP-OPD training before RL, finding that our two-stage framework achieves the highest scores among the methods evaluated. RP-OPD + RL shows limited signs of reward hacking on RubricHub Science, whereas the SFT + RL baseline increasingly receives high rewards for claims of rubric compliance without providing the required content. These findings support using rubrics to guide on-policy distillation before applying rubric-based RL.
Problem

Research questions and friction points this paper is trying to address.

Rubric-based reinforcement learning
Credit assignment
Reward hacking
On-policy distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rubric-based Reinforcement Learning
On-Policy Distillation
Token-level Supervision
Reward Hacking
Two-stage Training Framework
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.