Score
Design, implement, and evaluate multi-agent reinforcement learning systems that incorporate agent reputation or trust signals into rewards, policies, or critics—including training MADDPG-based models and trust/reputation-aware MADDPG variants. Build and analyze architectures that handle hybrid discrete–continuous action spaces and jointly optimize coordination objectives such as resource allocation and motion/control trajectories while measuring how reputation information alters learning and emergent behaviors.
This work addresses the computational bottleneck in multi-agent reinforcement learning (MARL) arising from centralized training in large-scale continuous action spaces. We propose a fully decentralized distributed training framework grounded in a networked communication topology. Departing from global centralized critics, our approach employs a distributed actor-critic architecture coupled with a surrogate policy mechanism, enabling cooperative policy optimization using only local neighbor communication. The framework natively supports cooperative, competitive, and mixed-task scenarios. Experiments demonstrate that our method matches MADDPG’s performance across diverse benchmark tasks while substantially reducing per-step training cost; this efficiency advantage scales favorably with increasing agent count. Overall, it establishes a new paradigm for scalable, low-overhead decentralized MARL.
This work addresses the challenges of training instability and inefficient coordination in multi-agent reinforcement learning caused by environmental non-stationarity. To enhance collaborative performance, the authors propose an action inference mechanism that enables each agent to explicitly predict the behaviors of its teammates. Furthermore, they introduce— for the first time in multi-agent settings—importance sampling based on the geometric distribution into experience replay, effectively mitigating non-stationarity. Integrated within the MADDPG framework, the proposed approach demonstrates substantially improved learning stability, exploration efficiency, and team coordination on the Predator-Prey task from the PettingZoo benchmark, consistently outperforming the standard MADDPG algorithm.
In multi-agent reinforcement learning (MARL), insufficient identification and incentivization of cooperative behaviors hinder joint and individual reward optimization. Method: This paper proposes a collaboration-enhanced algorithm built upon MADDPG, featuring a learnable dynamic cooperation identification module and a cooperative reward shaping mechanism that explicitly models and positively reinforces effective inter-agent collaboration. The approach is implemented within the PettingZoo benchmark suite, integrating multi-agent policy optimization, reward mechanism design, and a distributed training framework. Contribution/Results: Experiments across diverse cooperative and competitive tasks demonstrate significant improvements—average team cumulative reward increases by 23.6%—along with enhanced individual policy stability, higher collaboration efficiency, and faster convergence compared to baseline MADDPG. These results validate the effectiveness and generalizability of cooperation-aware modeling in advancing MARL performance.
This work addresses the scalability bottleneck in multi-agent reinforcement learning caused by the linear growth of input dimensionality with the number of agents in centralized critics. To overcome this limitation, the authors propose MADDPG-K, which introduces a k-nearest-neighbor mechanism based on Euclidean distance into the MADDPG framework for the first time. This allows each critic to observe only a fixed number of neighboring agents, thereby maintaining constant input dimensionality. While preserving the centralized training with decentralized execution paradigm, MADDPG-K significantly reduces computational complexity. Experimental results demonstrate that MADDPG-K achieves performance comparable to or better than MADDPG in multi-particle cooperative and competitive tasks, with faster convergence and markedly improved computational efficiency and scalability as the number of agents increases.
In multi-agent reinforcement learning (MARL), the exponential growth of the joint state-action space impedes efficient exploration, while high-quality collaborative expert demonstrations are often impractical to obtain. Method: This paper proposes the Personalized Expert Demonstration (PED) paradigm, wherein each agent receives only single-agent demonstrations focused on its individual objective—eliminating reliance on cooperative demonstrations. We introduce a novel dual-discriminator architecture: a behavior-alignment discriminator enforcing policy fidelity to individual demonstrations, and a goal-oriented discriminator incentivizing task completion. The framework integrates inverse reinforcement learning with adversarial reward shaping and supports both discrete and continuous action spaces. Contribution/Results: PED significantly outperforms state-of-the-art methods across diverse coordination benchmarks; exhibits strong robustness to suboptimal demonstrations; and successfully transfers non-cooperative policies to achieve rapid convergence in StarCraft II micromanagement tasks.
This study addresses the limitations of historical reputation in multi-agent debate, where it poorly predicts behavior on novel tasks and remains vulnerable to adversarial attacks. To overcome these challenges, this work proposes MiniRep, a system that innovatively aggregates real-time task behavior with long-term reputation dynamically. Furthermore, it introduces a taxonomy-based defense mechanism to prevent homogeneous response groups from dominating decision-making, thereby mitigating compound threats involving reputation manipulation and software mutation. Evaluated in a 10-agent heterogeneous setting on the MATH benchmark, MiniRep significantly outperforms conventional baseline methods across all 28 attack configurations. These results demonstrate that the proposed approach effectively enhances both the robustness and security of multi-agent collaboration under complex adversarial conditions.
研究通过引入基于声誉调整的强化学习方法,在空间囚徒困境游戏中促进合作行为的出现,展示了声誉不仅提供直接激励还重塑了社会信息环境。
This work addresses the challenge of Byzantine agents—those exhibiting faulty or malicious behavior—that undermine cooperative decision-making in multi-agent reinforcement learning (MARL) communication systems, a problem inadequately tackled by existing approaches. The paper proposes BARD-MARL, a novel post-hoc diagnostic framework that uniquely integrates policy graph trajectories with Bayesian trust modeling grounded in BayesG. By jointly analyzing inter-agent policy dependencies and communication reliability, BARD-MARL enables precise identification of Byzantine nodes. Evaluated in SUMO traffic signal control scenarios with 25 to 100 agents under diverse adversarial conditions—including observation flipping and coordinated attacks—the method achieves an AUC-ROC of up to 0.982, substantially outperforming baseline methods. These results demonstrate its effectiveness, scalability, and the critical advantage of decoupling coordination, detection, and mitigation mechanisms.
This work addresses the challenge of designing a global reward function in heterogeneous multi-agent reinforcement learning that effectively accommodates diverse agent objectives. To this end, the authors propose MAGPIE, a method that leverages expert preference signals to construct individual reward models for each agent without requiring a predefined global reward. A monotonic aggregation mechanism is introduced to combine these agent-specific rewards into a unified global objective for policy optimization. Theoretically, this approach uniquely bridges agent-specific preference modeling with Nash equilibrium optimization, proving that decentralized preference learning converges to a Nash equilibrium. Empirical results demonstrate that MAGPIE achieves performance on par with handcrafted reward functions across standard multi-agent benchmarks and sequential production-line scenarios, validating its effectiveness in eliminating the need for precise reward engineering.
This study investigates the scalability of single-agent parameterized action reinforcement learning algorithms in multi-agent environments. Addressing the limitations of the conventional centralized training with decentralized execution (CTDE) paradigm, the work proposes a fully decentralized training architecture wherein each agent independently maintains its own policy and value networks while sharing a common experience replay buffer. Within this framework, the performance of MAGAC, MASAC, and MATQC is systematically evaluated across varying numbers of agents, with statistical significance assessed via ANOVA and Tukey’s HSD test. Experimental results demonstrate that MAGAC achieves substantial performance gains in multi-agent settings, whereas MASAC and MATQC exhibit limited improvements. Moreover, beyond five agents, performance gains diminish marginally while computational costs rise sharply, revealing a critical trade-off between scalability and efficiency.