After Cooperation Is Learned: Gradient Routing and Optimizer-Dependent Maintenance in Multi-Agent Reinforcement Learning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the instability of maintaining learned cooperative policies during continual training in multi-agent reinforcement learning. Building upon an Actor-Critic architecture, the authors formulate cooperation maintenance as a right-censored event-time problem. By isolating value gradient routing from critic existence and employing matched warm-start experiments with gradient auditing, they systematically evaluate how shared feature updates under different optimizers affect policy stability in MinEx and CleanUp-lite environments. The findings reveal that high reward scaling exacerbates the maintenance sensitivity of direct gradient routing. Furthermore, blocking this pathway or removing the critic preserves stability, demonstrating that the risk originates from the gradient routing mechanism rather than the critic itself. This work provides a theoretical foundation for the stable maintenance of cooperative strategies in multi-agent systems.
📝 Abstract
Cooperative MARL is commonly evaluated through cooperation discovery from random initialization, leaving open whether continued optimization can destabilize learned cooperation. Actor-critic comparisons can also conflate critic presence with value gradients entering shared actor representations. We study cooperation maintenance, defined as the survival of a behaviorally verified cooperative policy under continued training. We formulate maintenance as a right-censored event-time problem and compare matched warm starts: X0 allows value loss gradients to update shared actor features, X1 retains the critic while blocking those gradients, and X5 removes the learned critic as a critic-free reference. This isolates direct value-gradient access while controlling initialization, critic computation, and evaluation. Positive reward scaling preserves strategic preferences and equilibria while perturbing learning dynamics. Gradient audits confirm the intended routing pathways, and frozen-policy torso perturbations probe whether route-induced updates align with local cooperation boundaries. In confirmatory MinEx and CleanUp-lite experiments, higher scales selectively increase maintenance sensitivity in X0; X1 remains near the censoring ceiling, and X5 has no confirmed events in the tested settings. In CleanUp-lite, route-by-scale displacement is associated with reduced local cooperation margins; MinEx shows a weaker, optimizer-dependent effect. These results identify a conditional, scale-sensitive maintenance risk associated with direct value-gradient routing rather than a universal failure of critics.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Reinforcement Learning
Cooperation Maintenance
Gradient Routing
Actor-Critic
Value Gradients
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cooperation Maintenance
Gradient Routing
Right-Censored Event-Time
Multi-Agent Reinforcement Learning
Positive Reward Scaling
🔎 Similar Papers
No similar papers found.