An Improved Multi-Agent Algorithm for Cooperative and Competitive Environments by Identifying and Encouraging Cooperation among Agents

📅 2025-08-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In multi-agent reinforcement learning (MARL), insufficient identification and incentivization of cooperative behaviors hinder joint and individual reward optimization. Method: This paper proposes a collaboration-enhanced algorithm built upon MADDPG, featuring a learnable dynamic cooperation identification module and a cooperative reward shaping mechanism that explicitly models and positively reinforces effective inter-agent collaboration. The approach is implemented within the PettingZoo benchmark suite, integrating multi-agent policy optimization, reward mechanism design, and a distributed training framework. Contribution/Results: Experiments across diverse cooperative and competitive tasks demonstrate significant improvements—average team cumulative reward increases by 23.6%—along with enhanced individual policy stability, higher collaboration efficiency, and faster convergence compared to baseline MADDPG. These results validate the effectiveness and generalizability of cooperation-aware modeling in advancing MARL performance.

Technology Category

Multiagent Systems: TeamworkGame Theory and Economic Paradigms: Cooperative Game TheoryHumans and AI: Teamwork, Team formation

Application Category

User Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingResponsible Web: Machine-in-the-loop, human agency and autonomyGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
We propose an improved algorithm by identifying and encouraging cooperative behavior in multi-agent environments. First, we analyze the shortcomings of existing algorithms in addressing multi-agent reinforcement learning problems. Then, based on the existing algorithm MADDPG, we introduce a new parameter to increase the reward that an agent can obtain when cooperative behavior among agents is identified. Finally, we compare our improved algorithm with MADDPG in environments from PettingZoo. The results show that the new algorithm helps agents achieve both higher team rewards and individual rewards.
Problem

Research questions and friction points this paper is trying to address.

Improving multi-agent reinforcement learning for cooperative environments
Addressing shortcomings in existing algorithms like MADDPG
Enhancing both team and individual rewards through cooperation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introducing new parameter to increase cooperative rewards
Improving MADDPG algorithm for multi-agent environments
Enhancing both team and individual agent rewards
J
Junjie Qi
Department of Industrial Engineering, Nanjing University, Nanjing, 210046, China
S
Siqi Mao
Department of Mathematics, Dalian University of Technology, Panjin, 124000, China
T
Tianyi Tan
Department of Mathematics, Dalian University of Technology, Panjin, 124000, China