🤖 AI Summary
This study addresses the instability of maintaining learned cooperative policies during continual training in multi-agent reinforcement learning. Building upon an Actor-Critic architecture, the authors formulate cooperation maintenance as a right-censored event-time problem. By isolating value gradient routing from critic existence and employing matched warm-start experiments with gradient auditing, they systematically evaluate how shared feature updates under different optimizers affect policy stability in MinEx and CleanUp-lite environments. The findings reveal that high reward scaling exacerbates the maintenance sensitivity of direct gradient routing. Furthermore, blocking this pathway or removing the critic preserves stability, demonstrating that the risk originates from the gradient routing mechanism rather than the critic itself. This work provides a theoretical foundation for the stable maintenance of cooperative strategies in multi-agent systems.
📝 Abstract
Cooperative MARL is commonly evaluated through cooperation discovery from random initialization, leaving open whether continued optimization can destabilize learned cooperation. Actor-critic comparisons can also conflate critic presence with value gradients entering shared actor representations. We study cooperation maintenance, defined as the survival of a behaviorally verified cooperative policy under continued training. We formulate maintenance as a right-censored event-time problem and compare matched warm starts: X0 allows value loss gradients to update shared actor features, X1 retains the critic while blocking those gradients, and X5 removes the learned critic as a critic-free reference. This isolates direct value-gradient access while controlling initialization, critic computation, and evaluation. Positive reward scaling preserves strategic preferences and equilibria while perturbing learning dynamics. Gradient audits confirm the intended routing pathways, and frozen-policy torso perturbations probe whether route-induced updates align with local cooperation boundaries. In confirmatory MinEx and CleanUp-lite experiments, higher scales selectively increase maintenance sensitivity in X0; X1 remains near the censoring ceiling, and X5 has no confirmed events in the tested settings. In CleanUp-lite, route-by-scale displacement is associated with reduced local cooperation margins; MinEx shows a weaker, optimizer-dependent effect. These results identify a conditional, scale-sensitive maintenance risk associated with direct value-gradient routing rather than a universal failure of critics.