Decomposing Communication Gain and Delay Cost Under Cross-Timestep Delays in Cooperative Multi-Agent Reinforcement Learning

📅 2026-04-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of degraded information timeliness and action misalignment in multi-agent collaboration caused by cross-timestep communication delays. It formalizes delayed communication as a Decentralized Communication Partially Observable Markov Game (DeComm-POMG) and introduces a Communication Gain and Delay Cost (CGDC) decomposition mechanism. Building upon this, the authors propose an adaptive communication strategy that triggers message exchange only when the net benefit is positive. They develop CDCMA, an actor-critic framework integrating future observation prediction, CGDC-guided attention, and dynamic message requesting to enhance information fusion efficiency. Experimental results demonstrate that the proposed method significantly outperforms baseline approaches across diverse delay settings in Cooperative Navigation, Predator-Prey, and SMAC tasks, achieving superior performance, robustness, and generalization capability.

Technology Category

Multiagent Systems: Agent CommunicationPlanning, Routing, and Scheduling: Planning with Markov Models (MDPs, POMDPs)Humans and AI: Human-Aware Planning and Behavior Prediction

Application Category

Economics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystemsResponsible Web: Machine-in-the-loop, human agency and autonomySearch and Retrieval-Augmented AI: Agentic search
📝 Abstract
Communication is essential for coordination in \emph{cooperative} multi-agent reinforcement learning under partial observability, yet \emph{cross-timestep} delays cause messages to arrive multiple timesteps after generation, inducing temporal misalignment and making information stale when consumed. We formalize this setting as a delayed-communication partially observable Markov game (DeComm-POMG) and decompose a message's effect into \emph{communication gain} and \emph{delay cost}, yielding the Communication Gain and Delay Cost (CGDC) metric. We further establish a value-loss bound showing that the degradation induced by delayed messages is upper-bounded by a discounted accumulation of an information gap between the action distributions induced by timely versus delayed messages. Guided by CGDC, we propose \textbf{CDCMA}, an actor--critic framework that requests messages only when predicted CGDC is positive, predicts future observations to reduce misalignment at consumption, and fuses delayed messages via CGDC-guided attention. Experiments on no-teammate-vision variants of Cooperative Navigation and Predator Prey, and on SMAC maps across multiple delay levels show consistent improvements in performance, robustness, and generalization, with ablations validating each component.
Problem

Research questions and friction points this paper is trying to address.

cross-timestep delays
cooperative multi-agent reinforcement learning
partial observability
communication delay
temporal misalignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Communication Gain
Delay Cost
Cross-Timestep Delay
Multi-Agent Reinforcement Learning
CGDC-Guided Attention
💼 Related Jobs
No related jobs found.
Z
Zihong Gao
The State Key Laboratory for Manufacturing Systems Engineering, School of Automation Science and Engineering, Xi'an Jiaotong University
H
Hongjian Liang
The State Key Laboratory for Manufacturing Systems Engineering, School of Automation Science and Engineering, Xi'an Jiaotong University
L
Lei Hao
The State Key Laboratory for Manufacturing Systems Engineering, School of Automation Science and Engineering, Xi'an Jiaotong University
L
Liangjun Ke
The State Key Laboratory for Manufacturing Systems Engineering, School of Automation Science and Engineering, Xi'an Jiaotong University