Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses optimization stagnation caused by sparse rewards and off-policy data format dependency in reinforcement learning for large models, proposing the RGPO framework. This method introduces a dynamic scaffolding mechanism that temporarily guides generation using ground-truth answers, feeding back only high-reward self-generated solutions into an unguided environment. This design enables reasoning training that integrates online reinforcement learning with adaptive prompting, applicable to both text-only and multimodal vision-language models. Experimental results demonstrate that RGPO significantly outperforms RLVR baselines. Furthermore, ablation studies confirm that adaptive guidance constitutes the primary source of performance gains, validating the framework's effectiveness and generality in mitigating reward sparsity while preserving exploration freedom.
📝 Abstract
On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is often limited by reward sparsity: when a model fails to discover correct trajectories for difficult problems, the optimization process receives little useful signal and may stagnate. Existing approaches mitigate this issue by incorporating off-policy demonstrations, expert traces, or model-generated solutions, but they typically require the auxiliary data to match the format of the reinforcement-learning task, often relying on rejection sampling from stronger models to obtain suitable training trajectories. We introduce Rationale-Guided Policy Optimization (RGPO), a framework that adaptively leverages ground-truth rationale information according to the model's current capability while preserving its freedom to explore. Rather than treating reference solutions as fixed imitation targets, RGPO uses them as temporary scaffolds: rationales help the model generate improved responses, after which only higher-reward, model-generated solutions are transferred back to the original unguided setting. This design allows training to exploit available ground-truth information without requiring off-policy data to follow the same format as the RL task. Across both language-only and vision-language reasoning settings, RGPO consistently improves performance over RLVR baselines, and ablation studies show that adaptive rationale guidance is a key contributor to these gains. These results suggest that RGPO offers a practical and general approach for reducing reward sparsity, stabilizing reinforcement learning, and improving reasoning performance in both text-only and multimodal models.
Problem

Research questions and friction points this paper is trying to address.

reward sparsity
reinforcement learning
reasoning
large language models
policy optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rationale-Guided Policy Optimization
Adaptive Rationale Scaffolding
Reward Sparsity
On-policy Reinforcement Learning
Vision-Language Reasoning
🔎 Similar Papers
No similar papers found.