decentralized multi-agent navigation

Designs, implements, and evaluates decentralized multi‑agent navigation systems and control policies that enable multiple agents to move through a shared environment (e.g., collision‑aware motion, cooperative coverage) while coordinating without centralized decision making; this includes developing training and algorithmic frameworks that use centralized training with decentralized execution (CTDE) such as MADDPG to learn and analyze multi‑agent control behavior.

decentralizedmulti-agentnavigation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Is Centralized Training with Decentralized Execution Framework Centralized Enough for MARL?

May 27, 2023
YZ
Yihe Zhou
🏛️ Zhejiang University | China Electric Power Research Institute

Existing CTDE frameworks permit access to global state during training but suffer from insufficient exploitation of inter-agent cooperation cues and inefficient joint policy exploration due to enforced policy independence. To address this, we propose Centralized Advice with Decentralized Pruning (CADP), a novel paradigm that introduces an explicit cross-agent advice mechanism to facilitate efficient collaborative learning during training, while integrating differentiable smooth model pruning to eliminate redundant parameters and enhance policy consistency—without compromising fully decentralized execution. Evaluated on StarCraft II micromanagement and Google Research Football benchmarks, CADP consistently outperforms state-of-the-art CTDE methods, achieving significant improvements in joint policy exploration efficiency and cooperative generalization. Our approach provides a principled framework for enhancing multi-agent coordination under the CTDE paradigm.

CTDE framework limits global cooperative information sharingInefficient joint-policy exploration in current CTDE methodsNeed for decentralized execution with enhanced centralized training

An Initial Introduction to Cooperative Multi-Agent Reinforcement Learning

May 10, 2024
CA
Christopher Amato
🏛️ Northeastern University

Cooperative multi-agent reinforcement learning (Cooperative MARL) suffers from conceptual ambiguity regarding fundamental paradigms—particularly the distinctions and applicability boundaries among centralized training with centralized execution (CTE), centralized training with decentralized execution (CTDE), and fully decentralized training and execution (DTE)—under the common setting of global reward sharing. Method: This work establishes a unified analytical framework to systematically characterize the design principles, intrinsic relationships, and evolutionary trajectories of major approaches, including value-decomposition methods (e.g., VDN, QMIX, QPLEX) and centralized-critic methods (e.g., MADDPG, COMA, MAPPO). Contribution: The analysis rigorously clarifies long-standing conceptual confusions in Cooperative MARL, yielding a structured cognitive map that supports principled algorithm selection, fair method comparison, and informed investigation of open challenges. The framework serves both as a pedagogical tool for teaching and a foundational reference for research advancement.

Compares centralized vs decentralized training and execution methodsIntroduces cooperative multi-agent reinforcement learning (MARL) conceptsReviews key MARL algorithms like VDN, QMIX, and MADDPG

This work addresses decentralized multi-agent navigation in cluttered environments, proposing the first joint optimization framework for agent policies and reconfigurable environmental layouts (e.g., obstacle placements). Methodologically, it employs model-free policy gradient reinforcement learning and introduces a two-stage alternating optimization algorithm that concurrently updates distributed agent policies and environmental structure. Theoretical analysis establishes convergence to local minima of a time-varying non-convex optimization problem. A key finding is that the optimized environment autonomously forms implicit, motion-decoupled guidance structures—enhancing behavioral coordination without explicit communication or centralized control. Experiments across diverse dense scenarios demonstrate consistent superiority over baselines in navigation success rate, throughput efficiency, and collision rate, empirically validating that environmental configuration optimization delivers substantial gains for multi-agent collaborative navigation.

Co-optimize agent policies and reconfigurable environments for navigationDecentralized multi-agent navigation in cluttered reconfigurable spacesModel-free learning to improve agent-environment performance synergy

Safe Multi-Agent Reinforcement Learning for Behavior-Based Cooperative Navigation

Dec 20, 2023
MD
Murad Dawood
🏛️ University of Bonn | Lamarr Institute for Machine Learning and Artificial Intelligence | Center for Robotics | Technical University of Munich

This work addresses safe collaborative navigation for multi-robot systems without individual reference trajectories. Methodologically, it proposes a behavior-driven safe multi-agent reinforcement learning framework that employs only the formation centroid as the navigation target—eliminating conventional per-robot path planners—and integrates model predictive control (MPC) as an online safety filter to explicitly guarantee collision-free operation during both training and deployment. To our knowledge, this is the first approach achieving provably safe collaborative navigation under the no-individual-reference setting. The MPC constraints not only accelerate policy convergence but also enable safe online deployment on real robots even in early training stages. Extensive simulations and real-world experiments demonstrate zero collisions, faster target arrival compared to baselines, and robust practical performance—validating both efficacy and deployability.

Achieving faster convergence and real-world deployment safetyPreventing collisions during multi-agent reinforcement learningSafe cooperative navigation without individual robot targets

Multi-UAV Collision Avoidance using Multi-Agent Reinforcement Learning with Counterfactual Credit Assignment

Apr 19, 2022
SH
Shuangyao Huang
🏛️ University of Otago | Xi’an Jiaotong-Liverpool University

Existing multi-agent reinforcement learning (MARL) approaches for cooperative collision avoidance among small-scale UAV swarms (≤3 agents) suffer from poor adaptability to continuous action spaces, high computational complexity, and excessive energy consumption. Method: We propose MACA, a centralized-training-with-decentralized-execution MARL algorithm featuring an actor-critic architecture and a novel marginalized state-action counterfactual baseline to address the credit assignment problem precisely. We further introduce MACAEnv—a physics-aware simulation environment that faithfully models UAV dynamics and inter-agent interaction constraints. Results: Experiments demonstrate that MACA achieves over 16% higher average reward than state-of-the-art MARL baselines; compared to conventional collision-avoidance methods, it reduces task failure rate by 90% and cuts response time by more than 99%. MACA exhibits strong robustness across diverse scenarios, significantly enhancing both flight safety and energy efficiency.

Decentralized collision avoidance for small UAV swarmsEnergy-efficient cooperation among UAVsOvercoming continuous action space challenges

Latest Papers

What's happening recently
View more

This study addresses the challenge of cooperative collision avoidance among multiple spacecraft under intermittent ground station communication constraints. The authors propose a semi-decentralized partially observable Markov decision process (SDec-POMDP) framework that explicitly incorporates ground station visibility into the multi-agent decision-making model for the first time. To solve for joint maneuver strategies, they design an approximate recursive short-horizon semi-decentralized A* algorithm (RS-SDA*). Operating solely within actual communication windows, this approach significantly reduces coordination synchronization events—by 28.5% compared to continuous coordination—while closely approximating the maneuver performance of centralized planning. Moreover, it satisfies safety distance constraints more consistently than heuristic rule-based methods and minimizes unnecessary orbital deviations.

collision avoidancecommunication constraintsintermittent communication

This work addresses cooperative multiagent control under communication uncertainty by proposing a semi-decentralized partially observable Markov decision process (SDec-POMDP) framework. It is the first to model communication actions as a stochastic temporal process and unifies Dec-POMDP and Multiagent POMDP through a probabilistic representation of communication history, thereby enabling flexible explicit communication mechanisms. Building on this formulation, the authors develop the Recursive Small-step Semi-Decentralized A* (RS-SDA*) algorithm to compute exact optimal policies. Empirical evaluation across multiple standard benchmarks and a maritime medical evacuation scenario demonstrates the efficacy of the approach, offering both a theoretical foundation and a practical toolkit for modeling and optimizing communication in multiagent systems.

communication uncertaintycooperative agentsmultiagent control

This work proposes a novel paradigm that integrates diffusion models with multi-agent reinforcement learning (MARL) to address the limitations of centralized and purely decentralized approaches in multi-robot coordination. Centralized planning suffers from poor scalability, while fully decentralized methods struggle to effectively model inter-agent interactions. The proposed method enables each robot to generate trajectories independently using single-agent data, while incorporating a centrally trained MARL value function to guide the reverse denoising process of the diffusion model via gradient-based refinement. This approach achieves interaction-aware coordinated planning without requiring joint modeling or retraining for varying numbers of robots. By leveraging exponential tilting for distribution adjustment, the method reduces agent interference from 55.4% to 41.8% in a four-robot maze navigation simulation, significantly enhancing coordination performance while maintaining strong scalability.

decentralized coordinationinter-agent interactionsmulti-robot motion planning

This work proposes a decentralized multi-agent reinforcement learning (MARL) approach to address the challenges of communication constraints, dynamic obstacles, and partial observability in GNSS-denied indoor environments for collaborative multi-drone exploration. Implemented on the high-fidelity Godot simulation platform, the method integrates LiDAR-based perception with local occupancy map sharing and models the problem as a networked distributed POMDP (ND-POMDP) to enable communication-aware cooperative exploration in continuous action spaces. By abandoning conventional reliance on discrete actions, centralized control, prior maps, and persistent connectivity, the approach introduces curriculum learning and a lightweight neural architecture, significantly enhancing training efficiency, robustness, and scalability. This provides a practical and efficient solution for real-world deployment of multi-drone systems.

collaborative explorationcommunication constraintsdecentralized decision-making

This work addresses the scalability, robustness, and generalization limitations in multi-agent reinforcement learning that arise from reliance on global state information—particularly the fragility observed under dynamic team compositions or environmental changes. To overcome these challenges, the authors propose a fully decentralized coordination framework that eschews all privileged centralized information, relying instead solely on local observations and peer-to-peer multi-hop communication for collaborative decision-making. The key innovations include a Distributed Graph Attention Network (D-GAT) for implicit global state inference and a novel Distributed Graph Attention MAPPO (DG-MAPPO) algorithm based on local policies and value functions. Experimental results demonstrate that the proposed method significantly outperforms state-of-the-art CTDE approaches across multiple benchmarks—including StarCraftII, Google Research Football, and Multi-Agent MuJoCo—and is effective for both homogeneous and heterogeneous agent teams.

centralized trainingconstrained communicationdecentralized execution

Hot Scholars

KO

Keisuke Okumura

University of Cambridge & National Institute of Advanced Industrial Science and Technology (AIST)
Multi-Agent Path PlanningMulti-Robot Coordination
AP

Amanda Prorok

Professor of Collective Intelligence and Robotics, University of Cambridge
RoboticsComputer Science
OS

Ocan Sankur

CNRS, Université de Rennes
Formal methods
BS

Bruno Sinopoli

Electrical and Systems Engineering, Washington University in
System TheoryControlCyber-Physical Systems
MN

Monica Nicoli

Associate Professor, Politecnico di Milano
signal processingwireless communicationlocalizationsmart mobility