🤖 AI Summary
This study addresses the structural mismatch between unordered teams and ordered inputs in multi-agent Transformer policies, where relying solely on permutation robustness evaluation can obscure action collapse. We propose an action collapse diagnostic metric that jointly assesses permutation consistency and action diversity to distinguish genuine robustness from homogeneous behaviors, thereby correcting biases inherent in conventional evaluations. Building upon Proximal Policy Optimization (PPO), self-attention mechanisms, and equivariant regularization, our analysis reveals that weak equivariant penalties effectively balance robustness with diversity. Furthermore, we demonstrate that different team sizes necessitate differentiated tuning of regularization weights. These findings offer a refined evaluation framework for multi-agent reinforcement learning systems employing Transformer architectures.
📝 Abstract
Transformer policies are attractive for multi-agent robot learning because self-attention can model interactions among agents. However, multi-agent teams are unordered, while transformers typically process agents as ordered token sequences. We study how this mismatch affects cooperative navigation policies under agent-order permutations. Our results show that low permutation error alone can be misleading: policies may appear robust simply because all agents choose the same action. We therefore evaluate policies using both permutation-consistency metrics and action-collapse diagnostics, including action diversity, same-action fraction, and maximum action frequency. A PPO-ID baseline yields non-collapsed behavior but remains order-sensitive, while strong equivariance regularization can still induce homogeneous behavior. A weak equivariance penalty improves the robustness while preserving more diverse actions for teams with \(N=3\) agents, whereas teams with \(N=4\) agents require substantially smaller regularization weights. These findings suggest that multi-agent transformer policies should be evaluated not only by return and permutation robustness, but also by whether they maintain non-collapsed, differentiated multi-agent behavior.