🤖 AI Summary
In multi-agent reinforcement learning (MARL), insufficient identification and incentivization of cooperative behaviors hinder joint and individual reward optimization. Method: This paper proposes a collaboration-enhanced algorithm built upon MADDPG, featuring a learnable dynamic cooperation identification module and a cooperative reward shaping mechanism that explicitly models and positively reinforces effective inter-agent collaboration. The approach is implemented within the PettingZoo benchmark suite, integrating multi-agent policy optimization, reward mechanism design, and a distributed training framework. Contribution/Results: Experiments across diverse cooperative and competitive tasks demonstrate significant improvements—average team cumulative reward increases by 23.6%—along with enhanced individual policy stability, higher collaboration efficiency, and faster convergence compared to baseline MADDPG. These results validate the effectiveness and generalizability of cooperation-aware modeling in advancing MARL performance.
📝 Abstract
We propose an improved algorithm by identifying and encouraging cooperative behavior in multi-agent environments. First, we analyze the shortcomings of existing algorithms in addressing multi-agent reinforcement learning problems. Then, based on the existing algorithm MADDPG, we introduce a new parameter to increase the reward that an agent can obtain when cooperative behavior among agents is identified. Finally, we compare our improved algorithm with MADDPG in environments from PettingZoo. The results show that the new algorithm helps agents achieve both higher team rewards and individual rewards.