Dynamic one-time delivery of critical data by small and sparse UAV swarms: a model problem for MARL scaling studies

📅 2025-12-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the single-shot, point-to-point delivery of critical packets in sparse swarms of small UAVs, proposing a decentralized multi-agent reinforcement learning (MARL) control framework. To systematically evaluate scalability, we introduce the first standardized family of dynamic deterministic games specifically designed for MARL scalability research. We further design a robust baseline policy integrating motion-envelope constraints with Dijkstra-based path planning, yielding an interpretable and reproducible performance lower bound. Experimental results show that mainstream MARL algorithms—such as MAPPO and QMix—achieve near-baseline performance in small-scale swarms. However, as agent count increases, severe training instability and policy degradation emerge, revealing—for the first time—an empirical scalability bottleneck in current MARL methods under sparse cooperative settings.

Technology Category

Multiagent Systems: Adversarial AgentsMachine Learning: Scalability of ML SystemsPlanning, Routing, and Scheduling: Planning with Markov Models (MDPs, POMDPs)

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomySearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesEconomics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystems
📝 Abstract
This work presents a conceptual study on the application of Multi-Agent Reinforcement Learning (MARL) for decentralized control of unmanned aerial vehicles to relay a critical data package to a known position. For this purpose, a family of deterministic games is introduced, designed for scaling studies for MARL. A robust baseline policy is proposed, which is based on restricting agent motion envelopes and applying Dijkstra's algorithm. Experimental results show that two off-the-shelf MARL algorithms perform competitively with the baseline for a small number of agents, but scalability issues arise as the number of agents increase.
Problem

Research questions and friction points this paper is trying to address.

Develops MARL for UAV swarm data delivery
Introduces deterministic games for scaling studies
Tests baseline and MARL scalability with agent count
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Reinforcement Learning for decentralized UAV control
Deterministic games designed for MARL scaling studies
Baseline policy using motion envelopes and Dijkstra's algorithm
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mika Persson
Saab AB, 112 76, Gothenburg, Sweden
J
Jonas Lidman
Swedish Defence Research Agency (FOI), 164 90, Stockholm, Sweden
J
Jacob Ljungberg
Chalmers University of Technology and the University of Gothenburg, Department of Mathematical Sciences, 412 58, Gothenburg, Sweden
S
Samuel Sandelius
Saab AB, 112 76, Gothenburg, Sweden
A
Adam Andersson
Chalmers University of Technology and the University of Gothenburg, Department of Mathematical Sciences, 412 58, Gothenburg, Sweden