Score
Designs and implements training systems for Multi-Agent Proximal Policy Optimization (MAPPO) under the centralized-training decentralized-execution (CTDE) paradigm, including actor networks, centralized critics, experience collection and replay, and optimization loops. Builds and analyzes training procedures, reward/advantage estimation, credit assignment, hyperparameter schedules and stabilization techniques to jointly optimize collective objectives (e.g., social welfare or throughput) while ensuring stable convergence and enabling decentralized policy execution at test time.
Existing CTDE methods fail to fully exploit the advantages of centralized training and lack theoretical guarantees. This paper proposes MAGPO, a novel CTDE framework that enables efficient cooperative exploration and decentralized execution via centralized guidance and decentralized policy alignment. Its core contributions are: (1) an autoregressive joint policy modeling mechanism that supports scalable multi-agent cooperative exploration; and (2) a policy alignment constraint coupled with monotonic policy gradient optimization, establishing—for the first time—theoretical guarantees of monotonic policy improvement in CTDE. Evaluated across six environments and 43 tasks, MAGPO consistently outperforms mainstream CTDE baselines and achieves performance on par with or superior to fully centralized approaches, demonstrating its effectiveness, generalizability, and practicality.
Existing CTDE frameworks permit access to global state during training but suffer from insufficient exploitation of inter-agent cooperation cues and inefficient joint policy exploration due to enforced policy independence. To address this, we propose Centralized Advice with Decentralized Pruning (CADP), a novel paradigm that introduces an explicit cross-agent advice mechanism to facilitate efficient collaborative learning during training, while integrating differentiable smooth model pruning to eliminate redundant parameters and enhance policy consistency—without compromising fully decentralized execution. Evaluated on StarCraft II micromanagement and Google Research Football benchmarks, CADP consistently outperforms state-of-the-art CTDE methods, achieving significant improvements in joint policy exploration efficiency and cooperative generalization. Our approach provides a principled framework for enhancing multi-agent coordination under the CTDE paradigm.
This work addresses the high variance in advantage estimation caused by non-stationary teammate policies in multi-agent reinforcement learning, which undermines the effectiveness of ratio-based trust-region methods such as MAPPO and MASPO. To mitigate this issue, the authors propose MARS, a novel policy optimization objective that replaces conventional additive ratio clipping or soft penalty mechanisms with a multiplicative symmetric geometric barrier within the centralized training with decentralized execution (CTDE) framework. This design imposes unbounded penalties on probability ratios approaching zero while preserving informative gradients, thereby preventing policy collapse and vanishing gradients. Empirical evaluation across 47 tasks spanning eight benchmark environments demonstrates that MARS consistently matches or outperforms existing methods, and ablation studies confirm the critical role of the symmetric geometric barrier in its performance gains.
Cooperative multi-agent reinforcement learning (Cooperative MARL) suffers from conceptual ambiguity regarding fundamental paradigms—particularly the distinctions and applicability boundaries among centralized training with centralized execution (CTE), centralized training with decentralized execution (CTDE), and fully decentralized training and execution (DTE)—under the common setting of global reward sharing. Method: This work establishes a unified analytical framework to systematically characterize the design principles, intrinsic relationships, and evolutionary trajectories of major approaches, including value-decomposition methods (e.g., VDN, QMIX, QPLEX) and centralized-critic methods (e.g., MADDPG, COMA, MAPPO). Contribution: The analysis rigorously clarifies long-standing conceptual confusions in Cooperative MARL, yielding a structured cognitive map that supports principled algorithm selection, fair method comparison, and informed investigation of open challenges. The framework serves both as a pedagogical tool for teaching and a foundational reference for research advancement.
This work addresses the challenge of efficiently computing joint policy gradients for global return maximization within the centralized training with decentralized execution (CTDE) framework. The authors propose a novel policy optimization method based on sequential joint decision-making, which provides the first exact decentralized decomposition of the joint policy gradient. By introducing an action belief mechanism to coordinate agent interactions and integrating sequential action commitments, decentralized critics, and individual score functions, each agent can perform independent updates while collectively implementing a complete joint gradient step. The approach does not rely on value factorization assumptions or converge to suboptimal equilibria. It achieves significant performance gains over strong baselines on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo benchmarks, with advantages amplifying as the number of agents scales up.
This work addresses the scalability, robustness, and generalization limitations in multi-agent reinforcement learning that arise from reliance on global state information—particularly the fragility observed under dynamic team compositions or environmental changes. To overcome these challenges, the authors propose a fully decentralized coordination framework that eschews all privileged centralized information, relying instead solely on local observations and peer-to-peer multi-hop communication for collaborative decision-making. The key innovations include a Distributed Graph Attention Network (D-GAT) for implicit global state inference and a novel Distributed Graph Attention MAPPO (DG-MAPPO) algorithm based on local policies and value functions. Experimental results demonstrate that the proposed method significantly outperforms state-of-the-art CTDE approaches across multiple benchmarks—including StarCraftII, Google Research Football, and Multi-Agent MuJoCo—and is effective for both homogeneous and heterogeneous agent teams.
This work addresses the challenge of policy convergence and target localization accuracy in multi-agent reinforcement learning under non-stationary, high-dimensional observations, where MAPPO struggles due to observation-induced non-stationarity. The authors propose ERPPO, a novel approach that uniquely integrates dynamic entropy regularization with estimation of observational ambiguity. Specifically, a Distributional Spatio-Temporal Ambiguity (DSA) learner quantifies environmental uncertainty, enabling adaptive switching between L1 and L2 regularization: L1 promotes exploration in high-ambiguity regions, while L2 stabilizes optimization in low-ambiguity regions. Evaluated in an AirSim maritime search-and-rescue simulation, ERPPO significantly improves target localization accuracy, effectively suppresses false detections under visual uncertainty, and achieves more efficient policy gradient updates.
This work addresses the issues of excessive advising, training instability, and performance degradation in decentralized multi-agent reinforcement learning caused by neglecting teacher-student compatibility. To this end, the authors propose a consensus-based communication and knowledge-sharing framework that constructs a consensus model from local observations via contrastive learning and integrates an action-scoring mechanism. Designed for seamless plug-in compatibility within the decentralized training with decentralized execution (DTDE) paradigm, the method enables agents to adaptively accept advice during action selection according to consensus constraints, thereby effectively balancing exploration and exploitation. Experimental results demonstrate that the proposed approach significantly improves coordination efficiency, learning speed, and final performance on benchmark environments including Google Research Football and StarCraft II, outperforming existing DTDE baselines.
Traditional centralized reinforcement learning struggles to support heterogeneous multi-agent systems, concurrent multi-task execution, and fault tolerance, thereby limiting the scalability of large language model agents in complex environments. This work proposes a decoupled distributed collective training architecture: trainable models reside on the server side to optimize GPU utilization, while diverse agents execute on arbitrary client devices, enabling flexible scaling. The framework supports heterogeneous multi-model reinforcement formulate, task isolation, fault-tolerant execution, and real-time code iteration during training. Integrated with an automated scientific research system, it facilitates long-term experiments without human intervention. Its context-tracking and timeline-merging mechanisms yield 1.5–10× training speedups and have successfully replicated human researchers’ exploration workflows in large-scale clusters, completing multi-day autonomous reinforcement learning experiments.