MATES: Learning Multi-Agent Interactions by Transforming Observations for Frozen Single-Agent Policies

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究提出MATES框架,通过转换多智能体观察来适应预训练的单智能体策略,以解决多智能体强化学习中同时获取任务能力和协调的问题。
📝 Abstract
Multi-agent reinforcement learning (MARL) commonly trains decentralized policies from scratch, requiring agents to acquire individual task competence and coordination simultaneously. Yet many multi-agent problems admit a compatible single-agent counterpart in which the underlying task can be learned in isolation. We introduce Multi-Agent Observation Transformation for Existing Single-Agent Policies (MATES), an input-side adaptation framework for tasks whose multi-agent observations preserve the solo-task information while exposing separately identifiable neighbor information. From multi-agent experience, MATES learns a small adapter that maps this observation into the format expected by a frozen single-agent policy, inducing actions suited to the shared environment without updating the single-agent policy itself. MATES leaves the pretrained policy's internal architecture unchanged and retains the objectives and update procedures of the underlying MARL algorithm. We evaluate MATES using both on- and off-policy algorithms on lifelong pathfinding, navigation, and cooperative discovery, spanning discrete and continuous observation and action spaces. Across all evaluated settings, MATES optimizes only 3.5-7.3% as many parameters as full-policy training while consistently outperforming MARL training from scratch. It approaches the performance of full fine-tuning, remains competitive overall with demonstration-based baselines, and retains strong task performance at team sizes not encountered during training. These results provide evidence that, under this observation structure, effective multi-agent behavior can be learned without modifying the policy that encodes individual competence.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Reinforcement Learning
Single-Agent Policies
Observation Transformation
Decentralized Policies
Task Competence
Innovation

Methods, ideas, or system contributions that make the work stand out.

MATES
multi-agent reinforcement learning
observation transformation
single-agent policy
adapter learning
🔎 Similar Papers
2023-12-04IEEE/RJS International Conference on Intelligent RObots and SystemsCitations: 0