Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of reward ambiguity and failure attribution arising from information constraints in fully decentralized multi-agent reinforcement learning by proposing the CASTLE framework. Specifically, this method offline pre-trains dual world models—capturing local dynamics and semantic social interactions—with frozen parameters. During the online phase, it leverages simulator rollbacks to generate counterfactual candidate action outcomes, providing prospective contextual guidance for independent PPO policies and transforming traditional post-hoc diagnosis into pre-action comparison. Experimental results demonstrate that the proposed framework significantly outperforms existing state-of-the-art baselines on benchmark tasks such as Tag, achieving performance improvements of up to 10.67 points.
📝 Abstract
Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information structure renders the conventional reward signal ambiguous. A poor return may result from an ineffective ego action, an incompatible teammate response, or an effective opponent response, yet scalar rewards alone do not reveal which explanation is responsible. We argue that agents can learn more effectively by prospectively comparing the consequences of candidate actions rather than diagnosing failures only from realized returns. We introduce CASTLE (Counterfactual Action-conditioned Semantic Tokens for Local Execution in Decentralized MARL), an offline-training, online-in-context guidance framework with two complementary world models. A Local Dynamics World Model, offline pre-trained over agents' local trajectories, summarizes the agent's local trajectory dynamics and partial observability, while a Semantic-Social World Model predicts compact short-horizon task and social consequences for each candidate ego action. The latter is trained from counterfactual simulator rollouts that expose plausible teammate and opponent responses to alternative actions taken from the same logged rollout state. During online learning and execution, both world models remain frozen and are queried by agents using only locally available information. Their prediction logits provide in-context guidance to an independent PPO policy. Across 30 matched seeds on Tag, Spread, and Adversary in the benchmark multi-particle environments, our proposed CASTLE achieves the highest mean final score among the evaluated methods, exceeding the strongest baseline on each task by 10.67, 6.46, and 0.33 normalized points, respectively.
Problem

Research questions and friction points this paper is trying to address.

Decentralized Multi-Agent Reinforcement Learning
Independent Learning
Reward Ambiguity
Partial Observability
Local Information
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Reinforcement Learning
Counterfactual Reasoning
World Models
Decentralized Learning
In-context Guidance