🤖 AI Summary
This work addresses the computational bottleneck in multi-agent reinforcement learning (MARL) arising from centralized training in large-scale continuous action spaces. We propose a fully decentralized distributed training framework grounded in a networked communication topology. Departing from global centralized critics, our approach employs a distributed actor-critic architecture coupled with a surrogate policy mechanism, enabling cooperative policy optimization using only local neighbor communication. The framework natively supports cooperative, competitive, and mixed-task scenarios. Experiments demonstrate that our method matches MADDPG’s performance across diverse benchmark tasks while substantially reducing per-step training cost; this efficiency advantage scales favorably with increasing agent count. Overall, it establishes a new paradigm for scalable, low-overhead decentralized MARL.
📝 Abstract
In this paper, we devise three actor-critic algorithms with decentralized training for multi-agent reinforcement learning in cooperative, adversarial, and mixed settings with continuous action spaces. To this goal, we adapt the MADDPG algorithm by applying a networked communication approach between agents. We introduce surrogate policies in order to decentralize the training while allowing for local communication during training. The decentralized algorithms achieve comparable results to the original MADDPG in empirical tests, while reducing computational cost. This is more pronounced with larger numbers of agents.