MAS-OPD: On-Policy Distillation for Multi-agent Systems

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of sparse reinforcement learning rewards and the difficulty of designing local rewards in multi-agent joint post-training. To this end, we propose a multi-agent joint training framework based on Online Policy Distillation (OPD). Methodologically, dense supervisory signals are leveraged to replace sparse rewards, thereby enhancing collaborative efficacy. Furthermore, we innovatively define "role advantage" to quantify agent specialization differences and introduce a privileged attribution mechanism to effectively mitigate cross-agent interaction conflicts. Experimental results demonstrate that the proposed method achieves state-of-the-art average performance on code and mathematics benchmarks, significantly improving both role differentiation clarity and collaborative capabilities among multiple agents.
📝 Abstract
Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competence or reliable collaboration, so post-training a MAS jointly is central. Most attempts use reinforcement learning, whose team-level reward leaves undetermined which step of which agent brought about the outcome, while local rewards need redesigning per task. On-policy distillation (OPD) gives token-level teacher supervision on trajectories the student samples, a denser signal needing no local reward, yet is underexplored for the interdependent agents of a MAS. Two difficulties arise: building complementary specialization from a judgement of which role a behavior belongs to while preserving the knowledge all roles need, and turning cross-agent collaborative information into supervision OPD can exploit. We present MAS-OPD, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged information. Extensive experiments on code and mathematics benchmarks show that MAS-OPD attains the highest mean score at both student scales and leads the agents to develop clearer role specialization and more effective collaborative behavior.
Problem

Research questions and friction points this paper is trying to address.

Multi-agent Systems
On-Policy Distillation
Role Specialization
Credit Assignment
Collaboration
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Multi-Agent Systems
Role-Advantage Specialization
Privileged Attribution
Post-training
🔎 Similar Papers