SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the sparse reward problem in Reinforcement Learning with Verifiable Rewards (RLVR) and the overconfidence of self-teachers in Online Policy Self-Distillation (OPSD), which inadvertently penalizes long trajectories. To this end, we propose the SIPO framework, which introduces a contrastive self-teacher mechanism that achieves dense, token-level credit assignment without external supervision. This mechanism provides effective learning signals even when all rollouts fail, thereby mitigating inherent biases. Furthermore, SIPO jointly optimizes reinforcement learning and online self-distillation. Experimental results demonstrate that SIPO significantly outperforms RLVR and OPSD baselines across multiple reasoning and code generation benchmarks, offering an efficient training paradigm for long chain-of-thought reasoning tasks.
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current policy, scores them with environment rewards, and constructs two teacher contexts for each rollout by pairing the reference answer with mistakes made within the group. The model then re-evaluates its own responses under both contexts, using the difference between the two teacher log-probabilities as token-level feedback, so that biases shared by both contexts are expected to largely cancel. The resulting objective yields a token-level advantage for every rollout: the reward still sets the main direction of each update while the self-teacher redistributes credit across tokens. Even in groups where every rollout fails and group-relative advantages vanish, SIPO still provides a learning signal. By preserving direct optimization of the task reward while providing dense, token-level feedback, this approach bridges reinforcement learning and on-policy self-distillation. Extensive experiments across multiple reasoning and code-generation benchmarks demonstrate that SIPO outperforms both RLVR and OPSD baselines without an external teacher or additional generation.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
On-Policy Self-Distillation
Token-level Credit Assignment
Large Language Models
Sparse Rewards
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Instructing Policy Optimization
On-Policy Self-Distillation
Contrastive Self-Teacher
Token-level Credit Assignment
Reinforcement Learning with Verifiable Rewards
🔎 Similar Papers