Scalable Neighborhood-Based Multi-Agent Actor-Critic

๐Ÿ“… 2026-04-20
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the scalability bottleneck in multi-agent reinforcement learning caused by the linear growth of input dimensionality with the number of agents in centralized critics. To overcome this limitation, the authors propose MADDPG-K, which introduces a k-nearest-neighbor mechanism based on Euclidean distance into the MADDPG framework for the first time. This allows each critic to observe only a fixed number of neighboring agents, thereby maintaining constant input dimensionality. While preserving the centralized training with decentralized execution paradigm, MADDPG-K significantly reduces computational complexity. Experimental results demonstrate that MADDPG-K achieves performance comparable to or better than MADDPG in multi-particle cooperative and competitive tasks, with faster convergence and markedly improved computational efficiency and scalability as the number of agents increases.

Technology Category

Multiagent Systems: Adversarial AgentsMachine Learning: Scalability of ML SystemsSearch and Optimization: Learning to Search

Application Category

Economics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystemsResponsible Web: Machine-in-the-loop, human agency and autonomySearch and Retrieval-Augmented AI: Efficiency and scalability of Web search engines
๐Ÿ“ Abstract
We propose MADDPG-K, a scalable extension to Multi-Agent Deep Deterministic Policy Gradient (MADDPG) that addresses the computational limitations of centralized critic approaches. Centralized critics, which condition on the observations and actions of all agents, have demonstrated significant performance gains in cooperative and competitive multi-agent settings. However, their critic networks grow linearly in input size with the number of agents, making them increasingly expensive to train at scale. MADDPG-K mitigates this by restricting each agent's critic to the $k$ closest agents under a chosen metric which in our case is Euclidean distance. This ensures a constant-size critic input regardless of the total agent count. We analyze the complexity of this approach, showing that the quadratic cost it retains arises from cheap scalar distance computations rather than the expensive neural network matrix multiplications that bottleneck standard MADDPG. We validate our method empirically across cooperative and adversarial environments from the Multi-Particle Environment suite, demonstrating competitive or superior performance compared to MADDPG, faster convergence in cooperative settings, and better runtime scaling as the number of agents grows. Our code is available at https://github.com/TimGop/MADDPG-K .
Problem

Research questions and friction points this paper is trying to address.

multi-agent reinforcement learning
centralized critic
scalability
computational complexity
MADDPG
Innovation

Methods, ideas, or system contributions that make the work stand out.

scalable multi-agent reinforcement learning
neighborhood-based critic
MADDPG-K
centralized critic
k-nearest neighbors
๐Ÿ”Ž Similar Papers
No similar papers found.
R
Rasmus Jensen
Department of Computer Science, ETH Zรผrich
T
Tim Goppelsroeder
Department of Computer Science, ETH Zรผrich