Score
Designs, builds, and evaluates algorithms, policies, and training frameworks for multiple interacting decision-making agents, covering centralized‑training/decentralized‑execution and fully decentralized control, cooperative/collaborative coordination, communication and value‑decomposition methods, and hierarchical policy structures (e.g., multi‑agent DQN, hierarchical MARL). Develops methods and metrics to handle nonstationarity, stability and convergence, online/adaptive learning, generalization to varying agent counts, safety or constraint satisfaction, and task allocation/coordination under partial observability.
This work systematically investigates three canonical interaction paradigms in multi-agent reinforcement learning (MARL): federated cooperation, decentralized collaboration, and non-cooperative games—corresponding respectively to centralized coordination, transient interaction, and incentive conflicts. Method: We propose the first unified taxonomy integrating federated learning principles into MARL, formally characterizing the theoretical boundaries and modeling assumptions of each paradigm’s topology. Leveraging tools from Markov decision processes, distributed optimization, and game theory, we conduct rigorous theoretical analysis and empirical evaluation. Contribution/Results: Our study clarifies fundamental trade-offs across convergence guarantees, communication efficiency, and equilibrium stability among paradigms, and identifies shared bottlenecks—including heterogeneity, non-stationarity, and incentive incompatibility—in existing approaches. The framework provides a structured conceptual foundation for MARL interaction modeling and informs future research directions toward practical deployment.
To address the four core challenges in multi-agent reinforcement learning (MARL)—non-stationarity, partial observability, large-scale scalability, and decentralized learning—this paper proposes a game-theoretic deep learning framework for cooperative learning. Methodologically, it unifies Nash equilibrium, evolutionary dynamics, correlated equilibrium, and adversarial dynamics into a single differentiable analytical paradigm, enabling gradient-based mapping from game-theoretic solutions to distributed policy updates. The framework integrates stochastic game modeling, projection-based policy-space gradient methods, population-level evolutionary differential equation approximations, and a decentralized actor-critic architecture. Empirically, on mixed cooperative-competitive benchmarks, it achieves a 42% improvement in convergence stability and a 37% gain in policy robustness. Moreover, it supports real-time Nash approximation for systems with up to one thousand agents. The approach bridges theoretical rigor—grounded in dynamic game theory—with engineering robustness, offering both formal guarantees and practical scalability.
This work addresses decentralized combinatorial optimization in dynamic multi-agent systems. We propose a hierarchical framework integrating reinforcement learning and collective learning: a high-level multi-agent reinforcement learning (MARL) module provides strategic guidance, while a low-level distributed collective learning mechanism enables cooperative decision-making. The architecture preserves agent autonomy while achieving scalability through action-space compression and minimal communication overhead, balancing long-term strategic planning, short-term collective performance, and environmental adaptability. Our key contribution is the first integration of MARL with decentralized collective learning to establish a scalable, Pareto-optimal evolutionary mechanism. Experiments in synthetic benchmarks and real-world smart-city applications—including energy self-management and drone swarm sensing—demonstrate significant improvements over standalone MARL or collective learning baselines, yielding superior optimization performance, system scalability, and robustness.
Cooperative multi-agent reinforcement learning (Cooperative MARL) suffers from conceptual ambiguity regarding fundamental paradigms—particularly the distinctions and applicability boundaries among centralized training with centralized execution (CTE), centralized training with decentralized execution (CTDE), and fully decentralized training and execution (DTE)—under the common setting of global reward sharing. Method: This work establishes a unified analytical framework to systematically characterize the design principles, intrinsic relationships, and evolutionary trajectories of major approaches, including value-decomposition methods (e.g., VDN, QMIX, QPLEX) and centralized-critic methods (e.g., MADDPG, COMA, MAPPO). Contribution: The analysis rigorously clarifies long-standing conceptual confusions in Cooperative MARL, yielding a structured cognitive map that supports principled algorithm selection, fair method comparison, and informed investigation of open challenges. The framework serves both as a pedagogical tool for teaching and a foundational reference for research advancement.
Existing CTDE frameworks permit access to global state during training but suffer from insufficient exploitation of inter-agent cooperation cues and inefficient joint policy exploration due to enforced policy independence. To address this, we propose Centralized Advice with Decentralized Pruning (CADP), a novel paradigm that introduces an explicit cross-agent advice mechanism to facilitate efficient collaborative learning during training, while integrating differentiable smooth model pruning to eliminate redundant parameters and enhance policy consistency—without compromising fully decentralized execution. Evaluated on StarCraft II micromanagement and Google Research Football benchmarks, CADP consistently outperforms state-of-the-art CTDE methods, achieving significant improvements in joint policy exploration efficiency and cooperative generalization. Our approach provides a principled framework for enhancing multi-agent coordination under the CTDE paradigm.
In decentralized execution for cooperative multi-agent reinforcement learning (MARL), mainstream decentralized policy gradient methods suffer from inherent suboptimality, preventing convergence to globally optimal policies. Method: We propose the Transformation-and-Distillation (TAD) framework, which equivalently reformulates a cooperative multi-agent MDP into a sequential single-agent MDP and employs policy distillation to recover decentralized execution. Contribution/Results: We theoretically prove that TAD guarantees learning of globally optimal policies in finite MDPs. Instantiating TAD with PPO, we develop TAD-PPO—incorporating MDP structural transformation, two-stage training, and value decomposition analysis. Empirical evaluation across diverse cooperative benchmarks demonstrates that TAD-PPO significantly outperforms state-of-the-art methods, achieving both theoretical global optimality guarantees and strong generalization capability.
This work addresses the inefficiencies in coordination and insufficient policy robustness in cooperative multi-agent reinforcement learning (MARL) that arise from relying solely on either local or global perspectives. To overcome this limitation, the paper proposes a Hierarchical Leader-Critic (HLC) architecture inspired by team organizational structures. HLC introduces, for the first time, a multi-level perspective mechanism into MARL, enabling synergistic learning of local and global information without explicit inter-agent communication by integrating high-level objectives with low-level execution. Coupled with a sequential training strategy, the proposed method significantly outperforms single-level baselines across multiple cooperative MARL benchmarks and demonstrates superior scalability and robustness as the number of agents and task complexity increase.
This work addresses the challenge of efficiently computing joint policy gradients for global return maximization within the centralized training with decentralized execution (CTDE) framework. The authors propose a novel policy optimization method based on sequential joint decision-making, which provides the first exact decentralized decomposition of the joint policy gradient. By introducing an action belief mechanism to coordinate agent interactions and integrating sequential action commitments, decentralized critics, and individual score functions, each agent can perform independent updates while collectively implementing a complete joint gradient step. The approach does not rely on value factorization assumptions or converge to suboptimal equilibria. It achieves significant performance gains over strong baselines on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo benchmarks, with advantages amplifying as the number of agents scales up.
This study addresses the single-shot, point-to-point delivery of critical packets in sparse swarms of small UAVs, proposing a decentralized multi-agent reinforcement learning (MARL) control framework. To systematically evaluate scalability, we introduce the first standardized family of dynamic deterministic games specifically designed for MARL scalability research. We further design a robust baseline policy integrating motion-envelope constraints with Dijkstra-based path planning, yielding an interpretable and reproducible performance lower bound. Experimental results show that mainstream MARL algorithms—such as MAPPO and QMix—achieve near-baseline performance in small-scale swarms. However, as agent count increases, severe training instability and policy degradation emerge, revealing—for the first time—an empirical scalability bottleneck in current MARL methods under sparse cooperative settings.
This work addresses the challenge of balancing global coordination and local execution in multi-agent reinforcement learning under asynchronous time scales. The authors propose a Coupled Hierarchical Multi-Agent System (CHMAS) that decomposes decision-making into centralized strategic planning and distributed tactical execution. A novel bidirectional feedback mechanism is introduced, wherein a coupling coefficient λ enables strategic objectives to dynamically adapt to tactical rewards, while an asynchronous update protocol mitigates environmental non-stationarity. The approach integrates a bilevel optimization framework, neighborhood-augmented distributed policies, and an analytically tractable additive approximation. Theoretical analysis establishes that the strategic layer achieves a convergence rate of 𝒪(log K/√K) after K updates. Empirical results demonstrate that the system converges stably and effectively learns spatially partitioned exploration strategies.
Existing collaborative multi-agent reinforcement learning (MARL) frameworks lack flexible, scalable, and end-to-end trainable architectures, often relying on fixed topologies or centralized training paradigms. Method: We propose Reinforcement Networks (RN), the first MARL framework that unifies agent systems as arbitrary directed acyclic graphs (DAGs), enabling modular, hierarchical, and graph-structured coordination. RN introduces a DAG-driven agent organization, end-to-end gradient propagation across the graph, graph-aware policy optimization, a novel collaboration-aware credit assignment algorithm, and LevelEnv—an environment abstraction for reproducible evaluation. Contribution/Results: Experiments demonstrate that RN consistently outperforms state-of-the-art baselines across diverse cooperative MARL benchmarks, achieving simultaneous improvements in task performance, scalability, and structural expressiveness. RN establishes a new paradigm for structured, scalable MARL grounded in principled graph-based representation and learning.