Fed-GRPO: Reward-Signal-Driven Federated Group Relative Policy Optimization

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of centralized training in reinforcement learning fine-tuning of large language models under privacy constraints by proposing the Federated GRPO framework. This framework introduces, for the first time, reward-statistics-based weighted aggregation, global calibration, and adaptive sparse communication mechanisms, enabling collaborative multi-node reasoning training without sharing raw data. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on mathematical reasoning tasks, closely approximating centralized training baselines. Furthermore, it reduces communication overhead by 32-fold and supports compression ratios of up to 621× with negligible accuracy degradation.
📝 Abstract
Large Language Models (LLMs) have shown strong reasoning capabilities when fine-tuned with reinforcement learning (RL), particularly through Group Relative Policy Optimization (GRPO). However, existing GRPO methods assume centralized access to training data, which may not hold in practice due to privacy or regulatory constraints. To this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication. Fed-GRPO contains three reward-signal-driven mechanisms: (i) \emph{signal-weighted aggregation} that weights clients by their reward standard deviation, prioritizing clients with stronger learning signals; (ii) \emph{global reward calibration} that re-weights per-prompt objectives based on the local-global reward gap, steering each client toward its relative weaknesses; and (iii) \emph{adaptive sparse communication} that allocates bandwidth based on the informativeness of each client's update. Extensive experiments on mathematical reasoning tasks demonstrate that Fed-GRPO achieves the best performance among all federated methods, clearly outperforms FedAvg and approaches centralized training performance, while losslessly reducing communication by $32\times$ and supporting up to $621\times$ compression under tight bandwidth budgets with only graceful accuracy degradation. Our code is available at https://github.com/HKU-HealthAI/Fed-GRPO.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Group Relative Policy Optimization
Federated Learning
Privacy Constraints
Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated Learning
Group Relative Policy Optimization
Large Language Models
Reward-Signal-Driven Mechanisms
Adaptive Sparse Communication
🔎 Similar Papers
No similar papers found.