π€ AI Summary
This work addresses decentralized dynamic task allocation for heterogeneous multi-agent systems (UAVs/UGVs) in 3D grid environments under communication constraints. We propose a Graph Neural Network (GNN)-driven Centralized Training with Decentralized Execution (CTDE) framework: (i) a GNN models agentβtask topological relationships; (ii) a lightweight message-passing mechanism and a conflict-aware sparse reward function are designed; (iii) Proximal Policy Optimization (PPO) is enhanced to handle non-stationarity; and (iv) conflict-free path planning integrates Reservation-based A* with R*. Experiments demonstrate a 92.5% collision-free rate, total flight time only 7.49% higher than the centralized Hungarian algorithm, and real-time scalability to 20 agents (2.8 s per timestep). The approach significantly improves adaptability to dynamic task arrivals and local robustness against partial observability and communication dropouts.
π Abstract
This paper addresses the challenge of decentralized task allocation within heterogeneous multi-agent systems operating under communication constraints. We introduce a novel framework that integrates graph neural networks (GNNs) with a centralized training and decentralized execution (CTDE) paradigm, further enhanced by a tailored Proximal Policy Optimization (PPO) algorithm for multi-agent deep reinforcement learning (MARL). Our approach enables unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs) to dynamically allocate tasks efficiently without necessitating central coordination in a 3D grid environment. The framework minimizes total travel time while simultaneously avoiding conflicts in task assignments. For the cost calculation and routing, we employ reservation-based A* and R* path planners. Experimental results revealed that our method achieves a high 92.5% conflict-free success rate, with only a 7.49% performance gap compared to the centralized Hungarian method, while outperforming the heuristic decentralized baseline based on greedy approach. Additionally, the framework exhibits scalability with up to 20 agents with allocation processing of 2.8 s and robustness in responding to dynamically generated tasks, underscoring its potential for real-world applications in complex multi-agent scenarios.