Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational overhead caused by multi-trajectory sampling and repeated evaluation in Group Relative Policy Optimization (GRPO), as well as the reduced learning efficiency due to low-quality trajectories. To this end, we propose FastRL, a plug-and-play reinforcement learning framework. It introduces a novel differential-aware advantage pruning algorithm that eliminates homogeneous trajectories to maximize gradient diversity, alongside an adaptive rolling sampling mechanism that dynamically adjusts the sampling scale to effectively balance exploration and efficiency. The framework is compatible with GRPO variants such as DAPO and GSPO. Experiments demonstrate that FastRL achieves an average 2.07× training speedup and a 1.64% improvement in visual reasoning accuracy on benchmarks including Geometry3K. The source code has been made publicly available.
📝 Abstract
Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning framework that simultaneously improves training efficiency and the effectiveness of policy learning. Specifically, 1) We introduce an advantage-aware pruning strategy to selectively preserve high-advantage trajectories while maximizing inter-trajectory gradient diversity. 2) Then, we design an adaptive rollout sampling mechanism to dynamically adjust the sampling scale across different training stages based on historical pruning distributions, balancing exploration adequacy and computational efficiency. Experiments demonstrate that FastRL can be seamlessly integrated into GRPO, DAPO, and GSPO variants, achieving an average 2.07$\times$ training speedup on Geometry3K and GeoQA8K-R1V, along with an approximately 1.64\% improvement in average accuracy on visual reasoning benchmarks. Source codes will be available at https://github.com/Nicozwy/FastRL.
Problem

Research questions and friction points this paper is trying to address.

Group Relative Policy Optimization
computational overhead
trajectory homogeneity
reinforcement learning
training efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Group Relative Policy Optimization
Advantage Pruning
Adaptive Rollout Sampling
Plug-and-Play Framework
Gradient Diversity
🔎 Similar Papers
No similar papers found.
J
Jiahua Yang
Guangdong Institute of Smart Education, Jinan University, Guangzhou, China
Zhiwei Yang
Zhiwei Yang
Guangzhou Institute of Technology, Xidian University, Guangzhou, China
Deep LearningComputer VisionAnomaly Detection
X
Xianpeng Zhang
OPPO AI Center, Shenzhen, China
D
Dongyu Chen
OPPO AI Center, Shenzhen, China
X
Xing Chen
Ragentile Intelligence Inc, Edmonton, Canada
T
Tianhuang Su
OPPO AI Center, Shenzhen, China
H
Haonan Lu
OPPO AI Center, Shenzhen, China
Quanlong Guan
Quanlong Guan
Jinan University
Multimodal LearningRepresentation learningRecommendation SystemAI in education
K
Kai Tang
OPPO AI Center, Shenzhen, China
C
Chuangchuang Wang
OPPO AI Center, Shenzhen, China