Score
Designs, implements, and evaluates decentralized multi‑agent navigation systems and control policies that enable multiple agents to move through a shared environment (e.g., collision‑aware motion, cooperative coverage) while coordinating without centralized decision making; this includes developing training and algorithmic frameworks that use centralized training with decentralized execution (CTDE) such as MADDPG to learn and analyze multi‑agent control behavior.
This work systematically investigates three canonical interaction paradigms in multi-agent reinforcement learning (MARL): federated cooperation, decentralized collaboration, and non-cooperative games—corresponding respectively to centralized coordination, transient interaction, and incentive conflicts. Method: We propose the first unified taxonomy integrating federated learning principles into MARL, formally characterizing the theoretical boundaries and modeling assumptions of each paradigm’s topology. Leveraging tools from Markov decision processes, distributed optimization, and game theory, we conduct rigorous theoretical analysis and empirical evaluation. Contribution/Results: Our study clarifies fundamental trade-offs across convergence guarantees, communication efficiency, and equilibrium stability among paradigms, and identifies shared bottlenecks—including heterogeneity, non-stationarity, and incentive incompatibility—in existing approaches. The framework provides a structured conceptual foundation for MARL interaction modeling and informs future research directions toward practical deployment.
Existing CTDE frameworks permit access to global state during training but suffer from insufficient exploitation of inter-agent cooperation cues and inefficient joint policy exploration due to enforced policy independence. To address this, we propose Centralized Advice with Decentralized Pruning (CADP), a novel paradigm that introduces an explicit cross-agent advice mechanism to facilitate efficient collaborative learning during training, while integrating differentiable smooth model pruning to eliminate redundant parameters and enhance policy consistency—without compromising fully decentralized execution. Evaluated on StarCraft II micromanagement and Google Research Football benchmarks, CADP consistently outperforms state-of-the-art CTDE methods, achieving significant improvements in joint policy exploration efficiency and cooperative generalization. Our approach provides a principled framework for enhancing multi-agent coordination under the CTDE paradigm.
Cooperative multi-agent reinforcement learning (Cooperative MARL) suffers from conceptual ambiguity regarding fundamental paradigms—particularly the distinctions and applicability boundaries among centralized training with centralized execution (CTE), centralized training with decentralized execution (CTDE), and fully decentralized training and execution (DTE)—under the common setting of global reward sharing. Method: This work establishes a unified analytical framework to systematically characterize the design principles, intrinsic relationships, and evolutionary trajectories of major approaches, including value-decomposition methods (e.g., VDN, QMIX, QPLEX) and centralized-critic methods (e.g., MADDPG, COMA, MAPPO). Contribution: The analysis rigorously clarifies long-standing conceptual confusions in Cooperative MARL, yielding a structured cognitive map that supports principled algorithm selection, fair method comparison, and informed investigation of open challenges. The framework serves both as a pedagogical tool for teaching and a foundational reference for research advancement.
This work addresses decentralized multi-agent navigation in cluttered environments, proposing the first joint optimization framework for agent policies and reconfigurable environmental layouts (e.g., obstacle placements). Methodologically, it employs model-free policy gradient reinforcement learning and introduces a two-stage alternating optimization algorithm that concurrently updates distributed agent policies and environmental structure. Theoretical analysis establishes convergence to local minima of a time-varying non-convex optimization problem. A key finding is that the optimized environment autonomously forms implicit, motion-decoupled guidance structures—enhancing behavioral coordination without explicit communication or centralized control. Experiments across diverse dense scenarios demonstrate consistent superiority over baselines in navigation success rate, throughput efficiency, and collision rate, empirically validating that environmental configuration optimization delivers substantial gains for multi-agent collaborative navigation.
This work addresses safe collaborative navigation for multi-robot systems without individual reference trajectories. Methodologically, it proposes a behavior-driven safe multi-agent reinforcement learning framework that employs only the formation centroid as the navigation target—eliminating conventional per-robot path planners—and integrates model predictive control (MPC) as an online safety filter to explicitly guarantee collision-free operation during both training and deployment. To our knowledge, this is the first approach achieving provably safe collaborative navigation under the no-individual-reference setting. The MPC constraints not only accelerate policy convergence but also enable safe online deployment on real robots even in early training stages. Extensive simulations and real-world experiments demonstrate zero collisions, faster target arrival compared to baselines, and robust practical performance—validating both efficacy and deployability.
Existing multi-agent reinforcement learning (MARL) approaches for cooperative collision avoidance among small-scale UAV swarms (≤3 agents) suffer from poor adaptability to continuous action spaces, high computational complexity, and excessive energy consumption. Method: We propose MACA, a centralized-training-with-decentralized-execution MARL algorithm featuring an actor-critic architecture and a novel marginalized state-action counterfactual baseline to address the credit assignment problem precisely. We further introduce MACAEnv—a physics-aware simulation environment that faithfully models UAV dynamics and inter-agent interaction constraints. Results: Experiments demonstrate that MACA achieves over 16% higher average reward than state-of-the-art MARL baselines; compared to conventional collision-avoidance methods, it reduces task failure rate by 90% and cuts response time by more than 99%. MACA exhibits strong robustness across diverse scenarios, significantly enhancing both flight safety and energy efficiency.
This study addresses the challenge of cooperative collision avoidance among multiple spacecraft under intermittent ground station communication constraints. The authors propose a semi-decentralized partially observable Markov decision process (SDec-POMDP) framework that explicitly incorporates ground station visibility into the multi-agent decision-making model for the first time. To solve for joint maneuver strategies, they design an approximate recursive short-horizon semi-decentralized A* algorithm (RS-SDA*). Operating solely within actual communication windows, this approach significantly reduces coordination synchronization events—by 28.5% compared to continuous coordination—while closely approximating the maneuver performance of centralized planning. Moreover, it satisfies safety distance constraints more consistently than heuristic rule-based methods and minimizes unnecessary orbital deviations.
This work addresses cooperative multiagent control under communication uncertainty by proposing a semi-decentralized partially observable Markov decision process (SDec-POMDP) framework. It is the first to model communication actions as a stochastic temporal process and unifies Dec-POMDP and Multiagent POMDP through a probabilistic representation of communication history, thereby enabling flexible explicit communication mechanisms. Building on this formulation, the authors develop the Recursive Small-step Semi-Decentralized A* (RS-SDA*) algorithm to compute exact optimal policies. Empirical evaluation across multiple standard benchmarks and a maritime medical evacuation scenario demonstrates the efficacy of the approach, offering both a theoretical foundation and a practical toolkit for modeling and optimizing communication in multiagent systems.
This work proposes a novel paradigm that integrates diffusion models with multi-agent reinforcement learning (MARL) to address the limitations of centralized and purely decentralized approaches in multi-robot coordination. Centralized planning suffers from poor scalability, while fully decentralized methods struggle to effectively model inter-agent interactions. The proposed method enables each robot to generate trajectories independently using single-agent data, while incorporating a centrally trained MARL value function to guide the reverse denoising process of the diffusion model via gradient-based refinement. This approach achieves interaction-aware coordinated planning without requiring joint modeling or retraining for varying numbers of robots. By leveraging exponential tilting for distribution adjustment, the method reduces agent interference from 55.4% to 41.8% in a four-robot maze navigation simulation, significantly enhancing coordination performance while maintaining strong scalability.
This work proposes a decentralized multi-agent reinforcement learning (MARL) approach to address the challenges of communication constraints, dynamic obstacles, and partial observability in GNSS-denied indoor environments for collaborative multi-drone exploration. Implemented on the high-fidelity Godot simulation platform, the method integrates LiDAR-based perception with local occupancy map sharing and models the problem as a networked distributed POMDP (ND-POMDP) to enable communication-aware cooperative exploration in continuous action spaces. By abandoning conventional reliance on discrete actions, centralized control, prior maps, and persistent connectivity, the approach introduces curriculum learning and a lightweight neural architecture, significantly enhancing training efficiency, robustness, and scalability. This provides a practical and efficient solution for real-world deployment of multi-drone systems.
This work addresses the scalability, robustness, and generalization limitations in multi-agent reinforcement learning that arise from reliance on global state information—particularly the fragility observed under dynamic team compositions or environmental changes. To overcome these challenges, the authors propose a fully decentralized coordination framework that eschews all privileged centralized information, relying instead solely on local observations and peer-to-peer multi-hop communication for collaborative decision-making. The key innovations include a Distributed Graph Attention Network (D-GAT) for implicit global state inference and a novel Distributed Graph Attention MAPPO (DG-MAPPO) algorithm based on local policies and value functions. Experimental results demonstrate that the proposed method significantly outperforms state-of-the-art CTDE approaches across multiple benchmarks—including StarCraftII, Google Research Football, and Multi-Agent MuJoCo—and is effective for both homogeneous and heterogeneous agent teams.