GRPO Training Dynamics for Small Language Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear training dynamics and poor reproducibility of Group Relative Policy Optimization (GRPO) in small language models by systematically investigating GRPO fine-tuning mechanisms for 1.5B- to 7B-parameter models within a single-node 8×A100 environment. By revealing tensor-level update dynamics and analyzing the impact of group size on convergence, we propose a joint optimization strategy combining mechanism-informed LoRA configuration with reward shaping. Experimental results demonstrate that this approach improves performance on mathematical benchmarks by approximately 80% while significantly enhancing scientific question answering and code reasoning capabilities. Ultimately, this work establishes a reliable paradigm for efficient reinforcement learning of small models in resource-constrained scenarios.
📝 Abstract
Group Relative Policy Optimization (GRPO) has emerged as a memory-efficient reinforcement fine-tuning (RFT) technique for reasoning-intensive tasks. How- ever, GRPO training dynamics on small language models (SLMs) remain poorly understood, limiting its reliable adoption and reproducibility in open and resource- constrained environments. In this work, we present a systematic study of GRPO fine-tuning for SLMs ranging from 1.5B to 7B parameters under a practical single- node 8xA100 compute budget. Our study spans multiple model families and reasoning domains, including mathematics, coding, and multiple-choice question answering (MCQ) in science. Across these settings, we analyze how group size affects policy convergence, training stability, and downstream benchmark per- formance. We further characterize tensor-level update dynamics during GRPO training and investigate whether the choice of LoRA target modules and layers can improve the performance of GRPO-tuned models. While our initial GRPO-tuned models outperform their base counterparts on approximately 80% of mathematical benchmark evaluations, they demonstrate limited capability on MCQ and code reasoning tasks. Guided by our mechanistic evaluations, we refined our LoRA and reward-shaping configurations to improve performance in latter domains. These findings provide practical guidance for GRPO training for SLMs.
Problem

Research questions and friction points this paper is trying to address.

Group Relative Policy Optimization
Small Language Models
Training Dynamics
Reinforcement Fine-Tuning
LoRA
Innovation

Methods, ideas, or system contributions that make the work stand out.

Group Relative Policy Optimization
Small Language Models
Reinforcement Fine-Tuning
LoRA
Reward Shaping
🔎 Similar Papers