🤖 AI Summary
Existing multi-objective multi-agent reinforcement learning (MOMARL) methods lack dedicated inner-loop actor-critic frameworks for continuous state-action spaces. Method: We propose MOMA-AC, the first inner-loop actor-critic architecture tailored for continuous MOMARL. It features a preference-conditioned multi-head actor and a centralized critic, enabling a single model to efficiently represent the global Pareto front; strategy learning and objective trade-off are decoupled to enhance scalability and training stability. Built upon TD3/DDPG enhancements and validated in multi-agent physics simulations, MOMA-AC targets cooperative locomotion tasks. Results: MOMA-AC significantly outperforms outer-loop and independent-training baselines, achieving substantial gains in hypervolume and expected utility. Crucially, its performance remains robust as the number of agents increases, demonstrating strong scalability and generalization in continuous MOMARL settings.
📝 Abstract
This paper addresses a critical gap in Multi-Objective Multi-Agent Reinforcement Learning (MOMARL) by introducing the first dedicated inner-loop actor-critic framework for continuous state and action spaces: Multi-Objective Multi-Agent Actor-Critic (MOMA-AC). Building on single-objective, single-agent algorithms, we instantiate this framework with Twin Delayed Deep Deterministic Policy Gradient (TD3) and Deep Deterministic Policy Gradient (DDPG), yielding MOMA-TD3 and MOMA-DDPG. The framework combines a multi-headed actor network, a centralised critic, and an objective preference-conditioning architecture, enabling a single neural network to encode the Pareto front of optimal trade-off policies for all agents across conflicting objectives in a continuous MOMARL setting. We also outline a natural test suite for continuous MOMARL by combining a pre-existing multi-agent single-objective physics simulator with its multi-objective single-agent counterpart. Evaluating cooperative locomotion tasks in this suite, we show that our framework achieves statistically significant improvements in expected utility and hypervolume relative to outer-loop and independent training baselines, while demonstrating stable scalability as the number of agents increases. These results establish our framework as a foundational step towards robust, scalable multi-objective policy learning in continuous multi-agent domains.