GRPODropout: Less is More for Online Reinforcement Learning Rollouts

πŸ“… 2026-10-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the degradation of exploration capability caused by policy entropy collapse in reinforcement learning for large language models. To this end, we propose an enhanced Group Relative Policy Optimization (GRPO) algorithm grounded in rollout-level theoretical analysis. By deriving a theoretical threshold, the method selectively discards a small fraction of high-probability positive-advantage rollouts and performs advantage recentering, thereby mitigating entropy collapse with minimal computational overhead. Both theoretically and empirically, this work demonstrates that β€œless is more”: simply optimizing how rollouts are utilized substantially improves learning efficiency. The proposed approach achieves higher accuracy and greater actor entropy with fewer samples, significantly outperforming the original GRPO baseline.
πŸ“ Abstract
Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such as reward modification and entropy/KL regularization, or through token-level reweighting. We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates. Under the same sampling budget, not all rollouts contribute positively to an update, and selectively excluding some can improve learning. To address this, we propose GRPODropout: before the standard update, we use a simple strategy that selectively removes a small number of high-probability positive-advantage rollouts and recenters the retained advantages. To motivate this design, we develop a rollout-level theoretical analysis that guides method design and threshold selection. The method changes only rollout usage, and adds negligible computational overhead. Experiments show higher accuracy than original GRPO and higher actor entropy while using fewer rollout samples for updates, illustrating "less is more." This work provides insight into RL rollout usage: removing some rollouts can improve performance. Code is available at https://github.com/hexuandeng/GRPODropout/.
Problem

Research questions and friction points this paper is trying to address.

policy entropy collapse
reinforcement learning
large language models
GRPO
rollout selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

GRPODropout
Policy Entropy Collapse
Online Reinforcement Learning
Rollout Selection
Large Language Models
H
Hexuan Deng
Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen); Beijing Zhongguancun Academy
Zihao Yan
Zihao Yan
University of Virginia
NanomaterialsCatalysisElectrochemistry
Xuebo Liu
Xuebo Liu
Associate Professor of Computer Science, Harbin Institute of Technology, Shenzhen
Large Language ModelsNatural Language ProcessingMachine Translation
S
Shuo Nie
Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen)
Y
Yue Wang
Beijing Zhongguancun Academy; XinzhuAI
C
Chen Wang
Beijing Zhongguancun Academy
Z
Zhaohua Zhang
Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen)
Tianwen Jiang
Tianwen Jiang
Harbin Institute of Technology
Knowledge GraphInformation ExtractionNatural Language Processing
Q
Qiuyong Xiao
Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen)
J
Jihong Zhang
Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen)
M
Min Zhang
Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen)