Save Your Saturated Data: Learning Beyond Reward Saturation in Group-Based RL

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of reward saturation in training data caused by the enhanced capabilities of large language models, which leads to the vanishing of relative signals in Group Relative Policy Optimization (GRPO). To mitigate this, we propose a multi-faceted intervention strategy encompassing data curation, generation, reward modeling, and advantage computation. The core innovation lies in introducing high-quality incorrect solutions as negative samples during the rollout phase, combined with an iterative data recycling technique that effectively recovers reinforcement learning signals from saturated data. Experimental results on Qwen3 models demonstrate performance improvements of 6.4% to 9.0% for GRPO, significantly outperforming conventional intervention methods. These findings validate the feasibility of efficiently reusing saturated data to sustain reinforcement learning progress.
📝 Abstract
Group-relative reinforcement learning (RL) relies on reward variation among sampled responses to estimate informative relative advantages. As language models become increasingly capable, existing training data can become reward-saturated: all sampled responses to the same problem might receive equally high rewards, where the group-relative learning signals vanish and leave previously useful data obsolete. In this work, we investigate whether useful learning signals can be recovered from such saturated data. We study interventions at four levels of group-policy RL pipelines---data, rollout, reward, and advantage---and conduct extensive RL training on saturated reasoning data only. While standard GRPO on saturated data would almost always yield near-0 advantages and near-noise signals, diverse interventions successfully recycle and repurpose such data: among the proposed strategies, interventions at rollout generation are consistently most effective: nudging the policy to generate ``high-quality'', incorrect solutions introduces rollouts with poor rewards into saturated groups as negative samples, which turns out to improve GRPO by 6.4% to 9.0% across Qwen3-1.7B and 4B. Other interventions such as increasing rollout temperature or adding auxiliary rewards can also restore non-zero advantages, but yield less consistent gains. Further analyses show that effective negative rollouts require informative negative trajectories, that the method remains effective alongside unsaturated data, and that it supports iterative recycling of newly saturated examples. While increasingly stronger LLMs would render more data as saturated, our results demonstrate that don't waste your saturated data: with the right strategies they can be recycled into useful RL training signals in an increasingly data-scarce world.
Problem

Research questions and friction points this paper is trying to address.

Reward Saturation
Group-Relative Reinforcement Learning
GRPO
Data Recycling
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward Saturation
Group-Relative Reinforcement Learning
GRPO
Rollout Intervention
Negative Sampling
🔎 Similar Papers