Score
Designs and implements decentralized reinforcement-learning agents and training pipelines that optimize decision-making policies across multiple clients or devices using Proximal Policy Optimization (PPO) algorithms and their federated variants. Builds and analyzes the federated update, client selection, communication and aggregation mechanisms and per-client policy adaptation, applying PPO-specific techniques such as clipped objectives, proximal updates or trust-region constraints to ensure stable, convergent training.
This work systematically investigates three canonical interaction paradigms in multi-agent reinforcement learning (MARL): federated cooperation, decentralized collaboration, and non-cooperative games—corresponding respectively to centralized coordination, transient interaction, and incentive conflicts. Method: We propose the first unified taxonomy integrating federated learning principles into MARL, formally characterizing the theoretical boundaries and modeling assumptions of each paradigm’s topology. Leveraging tools from Markov decision processes, distributed optimization, and game theory, we conduct rigorous theoretical analysis and empirical evaluation. Contribution/Results: Our study clarifies fundamental trade-offs across convergence guarantees, communication efficiency, and equilibrium stability among paradigms, and identifies shared bottlenecks—including heterogeneity, non-stationarity, and incentive incompatibility—in existing approaches. The framework provides a structured conceptual foundation for MARL interaction modeling and informs future research directions toward practical deployment.
In federated reinforcement learning (FRL), Proximal Policy Optimization (PPO) suffers from convergence degradation due to misalignment between local actor-critic update order and global model aggregation—particularly under data heterogeneity, where divergent critic estimates induce conflicting gradient directions. To address this, we propose FedRAC, the first FRL framework that reverses PPO’s standard update order to “actor before critic.” This design theoretically eliminates critic estimation bias across heterogeneous clients, yielding a convergence bound independent of data heterogeneity. FedRAC integrates policy-gradient-based actor updates, delayed critic synchronization, and rigorous convergence analysis. We evaluate it on three standard RL benchmarks and a highly heterogeneous SUMO autonomous driving task. Results show average cumulative reward improvements of 12.7%–23.4%, 1.8×–2.5× faster convergence, and significantly enhanced robustness and practicality.
This work addresses the longstanding challenge in reinforcement learning of reconciling theoretical stability with engineering efficiency in policy optimization. Methodologically, we propose an unconstrained first-order algorithm that modifies the PPO policy loss by introducing a tighter probability-ratio clipping mechanism; this implicitly enforces a KL-divergence constraint—without requiring second-order computations or explicit constraints (e.g., trust-region projections)—thereby guaranteeing monotonic improvement and convergence. Our key contribution is the first unified framework achieving TRPO-level theoretical robustness while retaining PPO-level implementation simplicity and computational efficiency. Empirical evaluation across multiple benchmark tasks demonstrates consistent and significant performance gains over standard PPO, particularly excelling in end-to-end training of large-scale neural networks: the method achieves higher sample efficiency, improved training stability, and superior generalization performance.
This work addresses the decentralized heterogeneous resource allocation problem in multi-agent systems. We propose LGTC-IPPO, a decentralized reinforcement learning framework featuring a novel dynamic clustering consensus mechanism that enables agents to autonomously form teams and jointly optimize decisions. The framework integrates distributed consensus algorithms, dynamic graph neural network–based clustering, and multi-timescale modeling of resource states, thereby facilitating local-consensus-driven resource reallocation—reducing reliance on global information and enabling, for the first time, real-time rescheduling of discharge-type decaying resources. Experiments across diverse scales and resource distributions demonstrate a 37% improvement in reward stability, a 2.1× increase in coordination efficiency over the IPPO baseline, and robust scalability to hundreds of agents.
Traditional evolutionary game models struggle to characterize the long-term cooperative evolution in spatial public goods games. Method: We propose a Proximal Policy Optimization (PPO) framework augmented with adversarial curriculum transfer—first introducing continuous-policy PPO into this domain and designing a two-stage adversarial curriculum training paradigm to model agent strategy evolution within dynamic spatial environments. Contribution/Results: Theoretically, we validate the “punishment promotes cooperation” hypothesis and uncover a novel mechanism wherein value-function propagation drives spatiotemporal payoff coordination. Empirically, our method triggers the cooperation phase transition earlier than standard PPO, Q-learning, and Fermi updating within critical enhancement-factor ranges; sustains stable cooperative equilibria; and demonstrates significantly enhanced robustness under challenging initial conditions—e.g., full-defection states.
Proximal Policy Optimization (PPO) suffers from high variance and high sample complexity, undermining training stability and efficiency. To address this, we propose HP3O—a novel PPO variant that introduces *trajectory recency modeling* into the PPO framework for the first time. Specifically, HP3O employs a FIFO trajectory replay buffer to reuse recent high-return trajectories and designs a mixed policy update mechanism under distributional shift constraints, jointly leveraging optimal and randomly sampled trajectories. We theoretically establish its monotonic policy improvement guarantee. Empirical evaluation on multiple continuous-control benchmark tasks demonstrates that HP3O significantly improves sample efficiency and training stability over standard PPO, A2C, and PPO-RND: it reduces policy gradient variance by 32% and consistently achieves superior final performance.
This work addresses the disconnect between trust region theory and the heuristic clipping objective in Proximal Policy Optimization (PPO) by introducing the Bounded Regularized Reinforcement Learning (BRRL) framework. The authors formulate a constrained regularized policy optimization problem, derive its closed-form optimal solution, and design the Bounded Policy Optimization (BPO) algorithm to minimize the advantage-weighted divergence between the parameterized policy and this solution. They further extend BPO to Generalized BPO (GBPO) for large language model (LLM) fine-tuning. This study provides the first theoretical justification for PPO-style methods, unifying trust region optimization with a cross-entropy perspective while guaranteeing monotonic policy improvement. Experiments demonstrate that BPO and GBPO consistently match or outperform PPO and GRPO across MuJoCo, Atari, IsaacLab, and LLM fine-tuning benchmarks in both stability and final performance.
This work addresses the vulnerability of decentralized multi-agent path planning to minor adversarial perturbations in observations, which can lead to erratic behaviors and system-wide congestion. To enhance resilience, the study introduces certified robustness into this domain for the first time, proposing two training strategies: adversarial training based on Adv-PPO and a fine-tuning approach incorporating the MACER stochastic smoothing regularizer. Both methods significantly improve robustness without altering the underlying network architecture or deployment pipeline. Experimental results on an 8×8 grid with four agents demonstrate that worst-case success rates increase from 2.5% to 59.2% with Adv-PPO alone and further to 77.5% ± 6.0% when combined with MACER, while incurring less than 1% performance degradation under unperturbed (clean) conditions.
This work addresses the trade-off between training stability and sample efficiency in deep reinforcement learning, where Proximal Policy Optimization (PPO) exhibits stable on-policy learning but suffers from low sample efficiency, while off-policy methods, though more efficient, are prone to estimation bias and variance. To reconcile these limitations, the paper proposes ExO-PPO, which constructs an expectation-based lower bound for generalized policy improvement, introduces a piecewise exponential clipping mechanism, and incorporates a multi-policy replay buffer to reuse historical trajectories. This approach preserves the stability of on-policy optimization while substantially enhancing sample efficiency. Experimental results demonstrate that ExO-PPO consistently outperforms standard PPO and its state-of-the-art variants across a range of tasks, achieving significant improvements in stability, sample efficiency, and overall performance.