🤖 AI Summary
This study addresses the challenges of coordination difficulty and opaque decision-making caused by partial observability in multi-agent reinforcement learning by proposing the ELV framework. Specifically, ELV employs a variational autoencoder to extract structured semantic concepts for constructing a global context, designs a dual-path attention mechanism to explicitly model decision logic, and introduces concept prediction error as an intrinsic reward to drive exploration. Experimental results demonstrate that ELV not only achieves competitive performance across multiple benchmark environments but also clearly reveals the agents' reasoning processes, effectively unifying high performance with interpretability.
📝 Abstract
Efficient cooperation is challenging due to the usual partial observability of each agent in multi-agent reinforcement learning. Recurrent networks encode local interaction histories, but their hidden representations provide limited insight into the information underlying individual decisions. To address these challenges, we propose a novel interpretable framework, called escaping local views (ELV), which introduces semantically structured latent concepts to render policy decisions transparent. Specifically, each agent extracts low-dimensional semantic concepts from its local observation and action-observation trajectory. These concepts are jointly encoded into a contextual latent variable via a variational autoencoder (VAE), which builds a bridge between local views and global semantics. To explicitly model the decision of each agent, we employ a dual-path attention mechanism in which one module estimates the salience of individual concepts relative to the global context, while the other captures higher-order cooperative patterns with pairwise concept interactions. Furthermore, we incorporate a concept prediction module that derives an intrinsic reward from next-concept prediction errors, which incentivizes agents to explore regions of semantic novelty. Experiments in multiple environments verify that ELV not only achieves competitive performance but also explicitly provides how agents reason about their decisions.