MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low sample efficiency of world models in multi-agent reinforcement learning caused by observation reconstruction. To this end, it proposes MA-JEPA, a framework that introduces the Joint Embedding Predictive Architecture (JEPA) into multi-agent world models for the first time. By predicting target representations rather than reconstructing pixels, the method effectively enhances the capture of control-relevant information. Technically, MA-JEPA employs a causal Transformer with categorical latent states to construct posterior and action-conditioned dynamics models, thereby enabling a model-based centralized training with decentralized execution (CTDE) paradigm. Evaluated on the SMAC benchmark, the proposed approach achieves or surpasses state-of-the-art baseline win rates across four maps, demonstrating its effectiveness in complex cooperative tasks.
📝 Abstract
World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal for multi-agent reinforcement learning. We introduce MA-JEPA, a stochastic world model that replaces observation reconstruction with prediction of target representations, enabling model-based multi-agent reinforcement learning with centralized training and decentralized execution. A categorical latent state and a causal Transformer are trained with posterior and action-conditioned dynamics prediction objectives and are then used for actor-critic learning from latent imagination. A training-only joint predictor conditions on all agents'local states and actions to predict each agent's next local observation embedding. These predictions are passed through the same local posterior used during real interaction with a centralized critic that is used only for value learning, with execution remaining decentralized. Our experiments show that this architecture performs strongly on SMAC, matching or exceeding the strongest reported comparator mean win rate on four of eight evaluated maps.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Reinforcement Learning
World Models
Joint-Embedding Predictive Architecture
Sample Efficiency
Decentralized Execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Reinforcement Learning
Joint-Embedding Predictive Architecture
World Models
Centralized Training Decentralized Execution
Causal Transformer
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Brandon Gary Kaplowitz
Department of Engineering Science, University of Oxford
O
Osaze James Obahor
Department of Engineering Science, University of Oxford
Christian Schroeder de Witt
Christian Schroeder de Witt
University of Oxford
Multi-agent LearningSecuritySafety