grpo reinforcement fine-tuning

Design and implement reinforcement-learning fine-tuning pipelines that apply the GRPO algorithm to a pretrained policy or model to optimize group-relative reward signals, including constructing reward estimators and the training loop. Build and analyze evaluation protocols and diagnostics to stabilize RL fine-tuning across tasks, improve task-specific performance and reasoning, and refine instruction-following behavior by tuning hyperparameters and algorithmic components.

grporeinforcementfine-tuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.43
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing LLM instruction-tuning algorithms—such as supervised fine-tuning (SFT), proximal policy optimization (PPO), and direct preference optimization (DPO)—are often explained with heavy reliance on prior knowledge, omit critical derivations, or remain overly abstract, resulting in high cognitive barriers and poor interpretability. Method: This paper systematically unifies mainstream reinforcement learning and preference optimization approaches under a concise, symbolically grounded derivation framework explicitly tailored to practical LLM training scenarios. Contribution/Results: We introduce GRAPE (Generalized Relative Advantage Policy Evolution), a novel paradigm for future preference learning designed to overcome fundamental limitations of current methods in objective design, training stability, and generalization. The framework provides a coherent, step-by-step exposition—from SFT through DPO—enhancing algorithmic intuition and theoretical transparency. It establishes a rigorous foundation for advancing preference-based LLM alignment and offers principled directions for subsequent research.

Explaining reinforcement learning algorithms for instruction tuningIntroducing new research directions with GRAPE frameworkProviding clear intuitive understanding of complex RL methods

To address the limitation of Vision-Language-Action (VLA) models—namely, their reliance on static trajectory datasets and inability to adapt to novel environments via online interactive feedback—this paper proposes a closed-loop, trajectory-level reinforcement learning framework. Methodologically, it introduces Groupwise Relative Policy Optimization (GRPO), the first trajectory-granular policy optimization mechanism that jointly leverages step-wise immediate rewards and trajectory-level success signals to enable more accurate advantage estimation and stable online training. The framework integrates vision-language-action joint modeling with dynamic advantage attribution, enabling online sampling and optimization of complete manipulation trajectories. Evaluated on the LIBERO-Object benchmark comprising ten robotic manipulation tasks, our approach significantly outperforms supervised fine-tuning (SFT) and multiple RL baselines, achieving a 12.7% absolute improvement in task completion rate while simultaneously enhancing policy robustness and cross-task generalization.

Enhancing policy robustness in diverse robotic manipulation tasksImproving VLA model fine-tuning via trajectory-wise RL optimizationOvercoming limitations of static dataset supervised fine-tuning

This study addresses the limited instruction-following and mathematical reasoning capabilities of lightweight language models (e.g., Qwen2.5-0.5B). We systematically investigate the efficacy of reinforcement learning (RL)-based fine-tuning for alignment. To this end, we conduct the first comparative evaluation—on small-scale models—of RLOO, DPO, and supervised fine-tuning (SFT) for instruction alignment. We further propose a novel inference-time strategy: “synthetic data augmentation + external verifier-guided Best-of-N reasoning”, enabling tool-augmented, verification-aware reasoning. Experimental results show that RLOO with DeBERTa-based reward modeling achieves optimal instruction alignment, while DPO demonstrates superior robustness. Crucially, mathematical reasoning accuracy improves significantly, validating the synergistic benefit of combining RL-based fine-tuning with external verification at inference time. Our work establishes a reproducible, computationally efficient technical pathway for aligning small language models and enhancing their reliability in complex reasoning tasks.

Comparison of SFT, DPO, and RLOO techniques for model alignmentEffectiveness of RL fine-tuning for instruction following and math reasoningImproving math accuracy via synthetic data and inference-time tools

This work addresses the high computational cost of replay-phase token generation in reinforcement learning fine-tuning of large language models, which typically requires generating full reasoning trajectories. The authors propose RPO, a plug-and-play algorithm that, for the first time, analyzes the contribution of individual segments within reasoning paths to final answer correctness and selectively optimizes only the critical suffix segments. By integrating an experience caching mechanism, RPO substantially reduces redundant token generation. The method is compatible with mainstream algorithms such as GRPO and DAPO, achieving up to 90% and 72% reductions in training time on 1.5B and 7B models, respectively, while cutting replay-phase token generation by approximately 95%—all without compromising performance relative to full-trajectory training.

computational overheadlarge language modelsreasoning trajectory

This work addresses the limitation of current large language models in mathematical reasoning due to the absence of an effective active reflection mechanism, which hinders their self-correction and reasoning capabilities. To overcome this, we propose a four-stage training framework that explicitly incorporates reflection rewards during training for the first time. Our approach synergistically optimizes cognitive and environmental interaction rewards through Group Relative Policy Optimization (GRPO), while integrating both accuracy and format-based reward signals. Full-parameter supervised fine-tuning is employed to enhance the model’s introspective capacity. Experimental results demonstrate state-of-the-art performance on mathematical reasoning benchmarks. Ablation studies confirm the critical role of reflection rewards and further show that full-parameter fine-tuning significantly outperforms parameter-efficient alternatives such as LoRA.

large language modelsmathematical reasoningreflection

Latest Papers

What's happening recently
View more

This work addresses the challenge of policy optimization stagnation in multi-turn tool-augmented reasoning, where sparse rewards and minimal intra-group reward variance hinder effective learning. To overcome this, the authors propose Reward-Conditioned Trajectory Policy (RCTP), which models exploration as a controllable guidance task by introducing discrete reward tokens. Within the GRPO framework, RCTP enhances intra-group trajectory diversity to improve advantage estimation. The approach integrates supervised fine-tuning (SFT) with GRPO-based reinforcement learning, leveraging special reward-target tokens for fine-tuning and generating high-quality trajectories through reward-conditioned rollouts. Evaluated on the BFCLv4 multi-turn benchmark, RCTP significantly outperforms existing baselines, with the Qwen-2.5-7B-Instruct model surpassing all closed-source API counterparts.

exploration efficiencygroup relative policy optimizationmulti-turn tool calling

Existing reinforcement learning fine-tuning algorithms—such as GRPO and DAPO—suffer from inconsistent design and formulation, hindering effective comparison and comprehension, particularly for non-expert users. To address this challenge, this work proposes the first interactive visualization tool that unifies multiple algorithms under a cohesive interface. By integrating three coordinated views—training overviews, step-level input-output inspection, and side-by-side algorithm comparison—the tool enables token-level tracking of policy optimization dynamics. Built upon frontend visualization technologies and training log parsing, it supports fine-grained representation of mainstream RL fine-tuning algorithms, substantially lowering the barrier to entry for educational demonstrations and algorithm selection. The tool is open-sourced and publicly accessible.

algorithm comparisonlarge language modelspolicy optimization

This work investigates whether gradient-free evolutionary strategies (ES) and gradient-based GRPO converge to geometrically similar solutions in the post-training of large language models. Despite achieving comparable task performance, we find that their parameter update directions are nearly orthogonal, with ES inducing substantially larger parameter changes and greater out-of-distribution behavioral drift. Through comprehensive analyses—including optimization trajectory inspection, KL divergence measurements, linear mode connectivity tests, and theoretical modeling—we demonstrate for the first time that, although no loss barrier separates their solutions, the two methods explore fundamentally distinct regions of the solution space. We further propose a unified theoretical framework to explain the geometric divergence between these optimization paradigms.

Evolution StrategiesForgettingGRPO

This work addresses the lack of theoretical understanding of GRPO training dynamics, which currently rely on empirical hyperparameter tuning. By leveraging first-principles reasoning, the authors formulate GRPO dynamics as a physical potential system influenced by an inertia term. Through dimensionality reduction, mean-field approximation, and softmax-bandit simplification, they derive a closed-form trajectory model that elucidates both the overdamped limit and oscillatory transition mechanisms. This model serves as a diagnostic tool capable of distinguishing among multiple failure modes. Empirical validation across three models and two group sizes demonstrates reward trajectory fits with R² ≥ 0.91, confirming group-size invariance and the predictability of stability thresholds.

failure modesGRPOhyperparameter

This work addresses a critical inefficiency in population-based reinforcement learning methods such as GRPO, where homogeneous rollout outcomes cause the advantage function to vanish, leading to absent gradient signals and wasted computation. To mitigate this issue, the authors propose AERO, a novel approach that dynamically avoids zero-advantage regions through adaptive control of rollout counts, a selective rejection mechanism, and Bayesian posterior estimation. AERO maintains effective policy optimization while substantially improving training efficiency. Empirical results demonstrate that, under an identical rollout budget, AERO reduces total training computation by approximately 48% and decreases per-step runtime by 45%, while matching or surpassing GRPO in performance on both Pass@8 and Avg@8 evaluation metrics.

Compute EfficiencyGroup Relative Policy OptimizationLLM Alignment

Hot Scholars

SS

Shuheng Shen

Ant Group
Machine LearningOptimizationPrivacy
YW

Yue Wen

University of Central Florida
ProstheticsRehabilitation roboticsMachine learningAdaptive control
WZ

Wangmeng Zuo

School of Computer Science and Technology, Harbin Institute of Technology
Computer VisionImage ProcessingGenerative AIDeep Learning
ZY

Zhengyuan Yang

Principal Researcher, Microsoft
Computer VisionMultimediaMultimodalPost-Training
ZX

Zhenyu Xu

Texas Tech University
Machine LearningProgram RepairLarge Language Model