federated policy optimization

Designs and implements decentralized reinforcement-learning agents and training pipelines that optimize decision-making policies across multiple clients or devices using Proximal Policy Optimization (PPO) algorithms and their federated variants. Builds and analyzes the federated update, client selection, communication and aggregation mechanisms and per-client policy adaptation, applying PPO-specific techniques such as clipped objectives, proximal updates or trust-region constraints to ensure stable, convergent training.

federatedpolicyoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning

Jun 02, 2025
ZX
Zhijie Xie
🏛️ The Hong Kong University of Science and Technology

In federated reinforcement learning (FRL), Proximal Policy Optimization (PPO) suffers from convergence degradation due to misalignment between local actor-critic update order and global model aggregation—particularly under data heterogeneity, where divergent critic estimates induce conflicting gradient directions. To address this, we propose FedRAC, the first FRL framework that reverses PPO’s standard update order to “actor before critic.” This design theoretically eliminates critic estimation bias across heterogeneous clients, yielding a convergence bound independent of data heterogeneity. FedRAC integrates policy-gradient-based actor updates, delayed critic synchronization, and rigorous convergence analysis. We evaluate it on three standard RL benchmarks and a highly heterogeneous SUMO autonomous driving task. Results show average cumulative reward improvements of 12.7%–23.4%, 1.8×–2.5× faster convergence, and significantly enhanced robustness and practicality.

Data heterogeneity causes divergent gradients in conventional PPO updatesPPO actor-critic update order challenges in Federated Reinforcement LearningReversed update order needed for global policy convergence in FRL

Simple Policy Optimization

Jan 29, 2024
ZX
Zhengpeng Xie
🏛️ The Hong Kong University of Science and Technology

This work addresses the longstanding challenge in reinforcement learning of reconciling theoretical stability with engineering efficiency in policy optimization. Methodologically, we propose an unconstrained first-order algorithm that modifies the PPO policy loss by introducing a tighter probability-ratio clipping mechanism; this implicitly enforces a KL-divergence constraint—without requiring second-order computations or explicit constraints (e.g., trust-region projections)—thereby guaranteeing monotonic improvement and convergence. Our key contribution is the first unified framework achieving TRPO-level theoretical robustness while retaining PPO-level implementation simplicity and computational efficiency. Empirical evaluation across multiple benchmark tasks demonstrates consistent and significant performance gains over standard PPO, particularly excelling in end-to-end training of large-scale neural networks: the method achieves higher sample efficiency, improved training stability, and superior generalization performance.

EfficiencyPerformanceStability

This work addresses the decentralized heterogeneous resource allocation problem in multi-agent systems. We propose LGTC-IPPO, a decentralized reinforcement learning framework featuring a novel dynamic clustering consensus mechanism that enables agents to autonomously form teams and jointly optimize decisions. The framework integrates distributed consensus algorithms, dynamic graph neural network–based clustering, and multi-timescale modeling of resource states, thereby facilitating local-consensus-driven resource reallocation—reducing reliance on global information and enabling, for the first time, real-time rescheduling of discharge-type decaying resources. Experiments across diverse scales and resource distributions demonstrate a 37% improvement in reward stability, a 2.1× increase in coordination efficiency over the IPPO baseline, and robust scalability to hundreds of agents.

Decentralized allocation of heterogeneous resources among agents.Dynamic cluster consensus for adaptive local sub-team formation.Scalable and robust performance with increasing agents and resources.

PPO-ACT: Proximal Policy Optimization with Adversarial Curriculum Transfer for Spatial Public Goods Games

May 07, 2025
ZY
Zhaoqilin Yang
🏛️ Guizhou University | Beijing Jiaotong University

Traditional evolutionary game models struggle to characterize the long-term cooperative evolution in spatial public goods games. Method: We propose a Proximal Policy Optimization (PPO) framework augmented with adversarial curriculum transfer—first introducing continuous-policy PPO into this domain and designing a two-stage adversarial curriculum training paradigm to model agent strategy evolution within dynamic spatial environments. Contribution/Results: Theoretically, we validate the “punishment promotes cooperation” hypothesis and uncover a novel mechanism wherein value-function propagation drives spatiotemporal payoff coordination. Empirically, our method triggers the cooperation phase transition earlier than standard PPO, Q-learning, and Fermi updating within critical enhancement-factor ranges; sustains stable cooperative equilibria; and demonstrates significantly enhanced robustness under challenging initial conditions—e.g., full-defection states.

Enhancing cooperation evolution mechanisms using adversarial curriculum transferModeling agent strategy optimization in dynamic spatial public goods gamesOvercoming limitations of traditional evolutionary game models in long-term decision-making

Enhancing PPO with Trajectory-Aware Hybrid Policies

Feb 21, 2025
QL
Qisai Liu
🏛️ Iowa State University

Proximal Policy Optimization (PPO) suffers from high variance and high sample complexity, undermining training stability and efficiency. To address this, we propose HP3O—a novel PPO variant that introduces *trajectory recency modeling* into the PPO framework for the first time. Specifically, HP3O employs a FIFO trajectory replay buffer to reuse recent high-return trajectories and designs a mixed policy update mechanism under distributional shift constraints, jointly leveraging optimal and randomly sampled trajectories. We theoretically establish its monotonic policy improvement guarantee. Empirical evaluation on multiple continuous-control benchmark tasks demonstrates that HP3O significantly improves sample efficiency and training stability over standard PPO, A2C, and PPO-RND: it reduces policy gradient variance by 32% and consistently achieves superior final performance.

Decreases sample complexityImproves data distribution driftReduces high variance in PPO

Latest Papers

What's happening recently
View more

This work addresses the disconnect between trust region theory and the heuristic clipping objective in Proximal Policy Optimization (PPO) by introducing the Bounded Regularized Reinforcement Learning (BRRL) framework. The authors formulate a constrained regularized policy optimization problem, derive its closed-form optimal solution, and design the Bounded Policy Optimization (BPO) algorithm to minimize the advantage-weighted divergence between the parameterized policy and this solution. They further extend BPO to Generalized BPO (GBPO) for large language model (LLM) fine-tuning. This study provides the first theoretical justification for PPO-style methods, unifying trust region optimization with a cross-entropy perspective while guaranteeing monotonic policy improvement. Experiments demonstrate that BPO and GBPO consistently match or outperform PPO and GRPO across MuJoCo, Atari, IsaacLab, and LLM fine-tuning benchmarks in both stability and final performance.

clipped objectivepolicy optimizationProximal Policy Optimization

This work addresses the vulnerability of decentralized multi-agent path planning to minor adversarial perturbations in observations, which can lead to erratic behaviors and system-wide congestion. To enhance resilience, the study introduces certified robustness into this domain for the first time, proposing two training strategies: adversarial training based on Adv-PPO and a fine-tuning approach incorporating the MACER stochastic smoothing regularizer. Both methods significantly improve robustness without altering the underlying network architecture or deployment pipeline. Experimental results on an 8×8 grid with four agents demonstrate that worst-case success rates increase from 2.5% to 59.2% with Adv-PPO alone and further to 77.5% ± 6.0% when combined with MACER, while incurring less than 1% performance degradation under unperturbed (clean) conditions.

Adversarial RobustnessDecentralized PlanningMulti-Agent Path Finding

This work addresses the trade-off between training stability and sample efficiency in deep reinforcement learning, where Proximal Policy Optimization (PPO) exhibits stable on-policy learning but suffers from low sample efficiency, while off-policy methods, though more efficient, are prone to estimation bias and variance. To reconcile these limitations, the paper proposes ExO-PPO, which constructs an expectation-based lower bound for generalized policy improvement, introduces a piecewise exponential clipping mechanism, and incorporates a multi-policy replay buffer to reuse historical trajectories. This approach preserves the stability of on-policy optimization while substantially enhancing sample efficiency. Experimental results demonstrate that ExO-PPO consistently outperforms standard PPO and its state-of-the-art variants across a range of tasks, achieving significant improvements in stability, sample efficiency, and overall performance.

off-policy learningon-policy trainingpolicy optimization

Hot Scholars

HD

Huaiyu Dai

Professor of Electrical and Computer Engineering, NC State University
CommunicationsSignal ProcessingNetworkingSecurity and Privacy
RZ

Ruichen Zhang

Nanyang Technological University
Next-generation NetworkingEdge IntelligenceAgentic AIReinforcement learning
WN

Wei Ni

FIEEE, AAIA Fellow, Senior Principal Scientist & Conjoint Professor, CSIRO/UNSW
6G security and privacyconnected and trusted intelligenceapplied AI/ML
TJ

Tao Jiang

Huazhong University of Science and Technology
Coding and ModulationMassive MIMOIntelligent symbiotic communicationSatellite-ground
JL

Jianqing Liu

Computer Science, NC State University
QuantumAINetworksSecurity