🤖 AI Summary
This study addresses the challenges of high sample complexity, absence of global information, and coordination difficulties in fully distributed multi-agent systems under networked discrete control. To this end, we propose a fully distributed multi-agent reinforcement learning algorithm incorporating safety constraints. The method introduces parallel neural network branches to process local observations, integrates a graph encoder with a state estimator to enhance local perception, and combines an Actor-Critic architecture with Control Barrier Functions to achieve built-in safety guarantees and policy constraints. Experimental results demonstrate that, compared to centralized policies, the proposed approach reduces average uncertainty by 26.3% while approximating the performance of computationally intensive attention-based algorithms with less than 1% error.
📝 Abstract
This paper develops a safe and fully decentralized multi-agent reinforcement learning (MARL) algorithm to solve a class of discrete-time control problems on networks, including the persistent monitoring problem. Fully decentralized control of agents, while offering numerous benefits, faces issues such as exponentially increasing sample complexity, lack of global information about the system, and challenges in coordinating between agents. To address these issues, this paper introduces a fully decentralized multi-agent reinforcement learning algorithm that integrates deep reinforcement learning with safety considerations. This method feeds a history of local observations of the network's state into two parallel neural-network branches: the graph encoder, which adds structural information and correlations among nodes, and a state estimator, which predicts the uncertainty at each node in the graph. Additionally, the result of feeding that input into an actor-critic network is passed through a discrete-time control barrier heuristic to reduce the likelihood that any node will be neglected. This approach enables teams of fully decentralized agents to solve challenging problems by increasing system awareness and incorporating built-in safety measures to prevent the adoption of potentially harmful control policies. Numerical results from a custom simulation environment demonstrate that the proposed algorithm achieves 26.3 percent lower average uncertainty than a centralized control policy and is within 1 percent of the uncertainty performance of a more computationally complex algorithm with added attention layers.