multi-agent reinforcement learning

Designs, builds, and evaluates algorithms, policies, and training frameworks for multiple interacting decision-making agents, covering centralized‑training/decentralized‑execution and fully decentralized control, cooperative/collaborative coordination, communication and value‑decomposition methods, and hierarchical policy structures (e.g., multi‑agent DQN, hierarchical MARL). Develops methods and metrics to handle nonstationarity, stability and convergence, online/adaptive learning, generalization to varying agent counts, safety or constraint satisfaction, and task allocation/coordination under partial observability.

multi-agentreinforcementlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.89
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$177K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

To address the four core challenges in multi-agent reinforcement learning (MARL)—non-stationarity, partial observability, large-scale scalability, and decentralized learning—this paper proposes a game-theoretic deep learning framework for cooperative learning. Methodologically, it unifies Nash equilibrium, evolutionary dynamics, correlated equilibrium, and adversarial dynamics into a single differentiable analytical paradigm, enabling gradient-based mapping from game-theoretic solutions to distributed policy updates. The framework integrates stochastic game modeling, projection-based policy-space gradient methods, population-level evolutionary differential equation approximations, and a decentralized actor-critic architecture. Empirically, on mixed cooperative-competitive benchmarks, it achieves a 42% improvement in convergence stability and a 37% gain in policy robustness. Moreover, it supports real-time Nash approximation for systems with up to one thousand agents. The approach bridges theoretical rigor—grounded in dynamic game theory—with engineering robustness, offering both formal guarantees and practical scalability.

Environmental InstabilityLimited Information and EfficiencyMulti-Robot Learning Systems

This work addresses decentralized combinatorial optimization in dynamic multi-agent systems. We propose a hierarchical framework integrating reinforcement learning and collective learning: a high-level multi-agent reinforcement learning (MARL) module provides strategic guidance, while a low-level distributed collective learning mechanism enables cooperative decision-making. The architecture preserves agent autonomy while achieving scalability through action-space compression and minimal communication overhead, balancing long-term strategic planning, short-term collective performance, and environmental adaptability. Our key contribution is the first integration of MARL with decentralized collective learning to establish a scalable, Pareto-optimal evolutionary mechanism. Experiments in synthetic benchmarks and real-world smart-city applications—including energy self-management and drone swarm sensing—demonstrate significant improvements over standalone MARL or collective learning baselines, yielding superior optimization performance, system scalability, and robustness.

Addresses decentralized combinatorial optimization in evolving multi-agent systemsReduces communication overhead while preserving agent autonomy and privacySolves exponential joint state-action space growth in multi-agent reinforcement learning

An Initial Introduction to Cooperative Multi-Agent Reinforcement Learning

May 10, 2024
CA
Christopher Amato
🏛️ Northeastern University

Cooperative multi-agent reinforcement learning (Cooperative MARL) suffers from conceptual ambiguity regarding fundamental paradigms—particularly the distinctions and applicability boundaries among centralized training with centralized execution (CTE), centralized training with decentralized execution (CTDE), and fully decentralized training and execution (DTE)—under the common setting of global reward sharing. Method: This work establishes a unified analytical framework to systematically characterize the design principles, intrinsic relationships, and evolutionary trajectories of major approaches, including value-decomposition methods (e.g., VDN, QMIX, QPLEX) and centralized-critic methods (e.g., MADDPG, COMA, MAPPO). Contribution: The analysis rigorously clarifies long-standing conceptual confusions in Cooperative MARL, yielding a structured cognitive map that supports principled algorithm selection, fair method comparison, and informed investigation of open challenges. The framework serves both as a pedagogical tool for teaching and a foundational reference for research advancement.

Compares centralized vs decentralized training and execution methodsIntroduces cooperative multi-agent reinforcement learning (MARL) conceptsReviews key MARL algorithms like VDN, QMIX, and MADDPG

Is Centralized Training with Decentralized Execution Framework Centralized Enough for MARL?

May 27, 2023
YZ
Yihe Zhou
🏛️ Zhejiang University | China Electric Power Research Institute

Existing CTDE frameworks permit access to global state during training but suffer from insufficient exploitation of inter-agent cooperation cues and inefficient joint policy exploration due to enforced policy independence. To address this, we propose Centralized Advice with Decentralized Pruning (CADP), a novel paradigm that introduces an explicit cross-agent advice mechanism to facilitate efficient collaborative learning during training, while integrating differentiable smooth model pruning to eliminate redundant parameters and enhance policy consistency—without compromising fully decentralized execution. Evaluated on StarCraft II micromanagement and Google Research Football benchmarks, CADP consistently outperforms state-of-the-art CTDE methods, achieving significant improvements in joint policy exploration efficiency and cooperative generalization. Our approach provides a principled framework for enhancing multi-agent coordination under the CTDE paradigm.

CTDE framework limits global cooperative information sharingInefficient joint-policy exploration in current CTDE methodsNeed for decentralized execution with enhanced centralized training

Towards Global Optimality in Cooperative MARL with the Transformation And Distillation Framework

Jul 12, 2022
JY
Jianing Ye
🏛️ Washington University in St. Louis | Tsinghua University

In decentralized execution for cooperative multi-agent reinforcement learning (MARL), mainstream decentralized policy gradient methods suffer from inherent suboptimality, preventing convergence to globally optimal policies. Method: We propose the Transformation-and-Distillation (TAD) framework, which equivalently reformulates a cooperative multi-agent MDP into a sequential single-agent MDP and employs policy distillation to recover decentralized execution. Contribution/Results: We theoretically prove that TAD guarantees learning of globally optimal policies in finite MDPs. Instantiating TAD with PPO, we develop TAD-PPO—incorporating MDP structural transformation, two-stage training, and value decomposition analysis. Empirical evaluation across diverse cooperative benchmarks demonstrates that TAD-PPO significantly outperforms state-of-the-art methods, achieving both theoretical global optimality guarantees and strong generalization capability.

Addresses suboptimality in decentralized MARL algorithmsEnsures decentralized execution with theoretical performance guaranteesProposes a framework for global optimality in cooperative tasks

Latest Papers

What's happening recently
View more

This work addresses the inefficiencies in coordination and insufficient policy robustness in cooperative multi-agent reinforcement learning (MARL) that arise from relying solely on either local or global perspectives. To overcome this limitation, the paper proposes a Hierarchical Leader-Critic (HLC) architecture inspired by team organizational structures. HLC introduces, for the first time, a multi-level perspective mechanism into MARL, enabling synergistic learning of local and global information without explicit inter-agent communication by integrating high-level objectives with low-level execution. Coupled with a sequential training strategy, the proposed method significantly outperforms single-level baselines across multiple cooperative MARL benchmarks and demonstrates superior scalability and robustness as the number of agents and task complexity increase.

Cooperative MARLHierarchical LearningMulti-Agent Reinforcement Learning

This work addresses the challenge of efficiently computing joint policy gradients for global return maximization within the centralized training with decentralized execution (CTDE) framework. The authors propose a novel policy optimization method based on sequential joint decision-making, which provides the first exact decentralized decomposition of the joint policy gradient. By introducing an action belief mechanism to coordinate agent interactions and integrating sequential action commitments, decentralized critics, and individual score functions, each agent can perform independent updates while collectively implementing a complete joint gradient step. The approach does not rely on value factorization assumptions or converge to suboptimal equilibria. It achieves significant performance gains over strong baselines on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo benchmarks, with advantages amplifying as the number of agents scales up.

Centralized Training with Decentralized ExecutionCooperative TasksJoint Policy Optimization

Dynamic one-time delivery of critical data by small and sparse UAV swarms: a model problem for MARL scaling studies

Dec 10, 2025
MP
Mika Persson
🏛️ Saab AB | Swedish Defence Research Agency (FOI) | Chalmers University of Technology | University of Gothenburg

This study addresses the single-shot, point-to-point delivery of critical packets in sparse swarms of small UAVs, proposing a decentralized multi-agent reinforcement learning (MARL) control framework. To systematically evaluate scalability, we introduce the first standardized family of dynamic deterministic games specifically designed for MARL scalability research. We further design a robust baseline policy integrating motion-envelope constraints with Dijkstra-based path planning, yielding an interpretable and reproducible performance lower bound. Experimental results show that mainstream MARL algorithms—such as MAPPO and QMix—achieve near-baseline performance in small-scale swarms. However, as agent count increases, severe training instability and policy degradation emerge, revealing—for the first time—an empirical scalability bottleneck in current MARL methods under sparse cooperative settings.

Develops MARL for UAV swarm data deliveryIntroduces deterministic games for scaling studiesTests baseline and MARL scalability with agent count

This work addresses the challenge of balancing global coordination and local execution in multi-agent reinforcement learning under asynchronous time scales. The authors propose a Coupled Hierarchical Multi-Agent System (CHMAS) that decomposes decision-making into centralized strategic planning and distributed tactical execution. A novel bidirectional feedback mechanism is introduced, wherein a coupling coefficient λ enables strategic objectives to dynamically adapt to tactical rewards, while an asynchronous update protocol mitigates environmental non-stationarity. The approach integrates a bilevel optimization framework, neighborhood-augmented distributed policies, and an analytically tractable additive approximation. Theoretical analysis establishes that the strategic layer achieves a convergence rate of 𝒪(log K/√K) after K updates. Empirical results demonstrate that the system converges stably and effectively learns spatially partitioned exploration strategies.

global coordinationhierarchical decision-makinglocal execution

Reinforcement Networks: novel framework for collaborative Multi-Agent Reinforcement Learning tasks

Dec 28, 2025
MK
Maksim Kryzhanovskiy
🏛️ Lomonosov Moscow State University

Existing collaborative multi-agent reinforcement learning (MARL) frameworks lack flexible, scalable, and end-to-end trainable architectures, often relying on fixed topologies or centralized training paradigms. Method: We propose Reinforcement Networks (RN), the first MARL framework that unifies agent systems as arbitrary directed acyclic graphs (DAGs), enabling modular, hierarchical, and graph-structured coordination. RN introduces a DAG-driven agent organization, end-to-end gradient propagation across the graph, graph-aware policy optimization, a novel collaboration-aware credit assignment algorithm, and LevelEnv—an environment abstraction for reproducible evaluation. Contribution/Results: Experiments demonstrate that RN consistently outperforms state-of-the-art baselines across diverse cooperative MARL benchmarks, achieving simultaneous improvements in task performance, scalability, and structural expressiveness. RN establishes a new paradigm for structured, scalable MARL grounded in principled graph-based representation and learning.

Develops a general framework for collaborative multi-agent reinforcement learningEnables flexible credit assignment and scalable coordination in arbitrary DAGsUnifies hierarchical, modular, and graph-structured views of MARL systems

Hot Scholars

GS

Guillaume Sartoretti

Assistant Professor, National University of Singapore (NUS), Mechanical Engineering Dpt
Multi-Agent SystemsRoboticsSwarm IntelligenceDistributed Control
NJ

Natasha Jaques

University of Washington, Google Research
Social reinforcement learningMachine learningdeep learningmulti-agent
RL

Rongpeng Li

Zhejiang University
Multi-Agent CommunicationsNetGPTMARLNetwork Slicing
YC

Yiqun Chen

Renmin University of China
Information RetrievalRetrieval-Augmented GenerationReinforcement LearningMulti-Agent Systems
PF

Pingyi Fan

Professor of Electronic Engineering, Tsinghua University
Wireless CommunicationsInformation TheoryComputer Science