🤖 AI Summary
This study addresses the challenge of credit assignment in open teams, where frequent agent turnover makes it difficult to disentangle action contributions from personnel changes. We propose TOCA, a novel value decomposition framework that decouples returns into action, turnover, and interaction effects to achieve orthogonal credit assignment. By integrating a permutation-invariant centralized critic, event tokens, and counterfactual signals, TOCA effectively filters out turnover noise while preserving robust action credits. Evaluated on dynamic diffusion benchmarks with high turnover rates, TOCA and its variants achieve superior returns, significantly outperforming existing baselines. This work establishes a new paradigm for open multi-agent collaboration by enabling precise credit attribution despite continuous team composition changes.
📝 Abstract
Open-team multi-agent reinforcement learning studies cooperative systems in which agents may join, leave, or be replaced during an episode. In such settings, the team return changes both because agents choose useful actions and because the active population itself changes. Standard centralized critics and shared advantages often mix these two effects into one scalar credit signal, allowing surviving agents to be rewarded or penalized for exogenous turnover events outside their control. We introduce turnover-orthogonal credit assignment (TOCA), a value decomposition for open teams that separates action effects, pure turnover effects, and action--turnover interactions. Under exogenous turnover, the event-conditioned value admits a centered decomposition whose event-conditioned baseline removes the pure turnover component while preserving credit for actions that make the team robust to future replacements. We instantiate this idea with a permutation-invariant centralized critic over variable-size agent sets and event tokens, and derive both a counterfactual per-agent credit signal and a softly weighted interaction variant, TOCA-$β$, for high-variance control environments. Controlled diagnostic experiments show that TOCA improves return over event-aware MAPPO-style critics and that removing interaction credit substantially hurts performance. In a replacement-only Dynamic Spread benchmark, TOCA-$β$ achieves the best mean return at high turnover rates and improves over its no-interaction ablation. These results suggest that explicitly separating turnover from action credit is a useful principle for robust learning in dynamic cooperative teams.