🤖 AI Summary
This study addresses the single-shot, point-to-point delivery of critical packets in sparse swarms of small UAVs, proposing a decentralized multi-agent reinforcement learning (MARL) control framework. To systematically evaluate scalability, we introduce the first standardized family of dynamic deterministic games specifically designed for MARL scalability research. We further design a robust baseline policy integrating motion-envelope constraints with Dijkstra-based path planning, yielding an interpretable and reproducible performance lower bound. Experimental results show that mainstream MARL algorithms—such as MAPPO and QMix—achieve near-baseline performance in small-scale swarms. However, as agent count increases, severe training instability and policy degradation emerge, revealing—for the first time—an empirical scalability bottleneck in current MARL methods under sparse cooperative settings.
📝 Abstract
This work presents a conceptual study on the application of Multi-Agent Reinforcement Learning (MARL) for decentralized control of unmanned aerial vehicles to relay a critical data package to a known position. For this purpose, a family of deterministic games is introduced, designed for scaling studies for MARL. A robust baseline policy is proposed, which is based on restricting agent motion envelopes and applying Dijkstra's algorithm. Experimental results show that two off-the-shelf MARL algorithms perform competitively with the baseline for a small number of agents, but scalability issues arise as the number of agents increase.