Improving LLM Collaboration via Multi-Agent Preference Learning

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of reward construction and the limited scalability of preference learning in large language model (LLM) multi-agent collaboration by proposing the MAPL framework. This work represents the first systematic integration of preference learning into multi-agent systems, optimizing collaborative strategies through iterative comparisons between decentralized and centralized paradigms. Specifically, it introduces Multi-Agent Reinforcement Learning from Human Feedback (MARLHF) and Multi-Agent Direct Preference Optimization (MADPO). Experimental results demonstrate that MAPL significantly enhances collaboration quality and efficiency across tasks such as writing and coding, achieving performance comparable to the upper bound of multi-agent reinforcement learning with fixed rewards. Ultimately, this research establishes a generalizable optimization paradigm for multi-agent collaboration.
📝 Abstract
Several works have explored multi-agent reinforcement learning (MARL) in LLM collaboration. However, constructing reliable rewards is difficult in practice, as complete and accurate metrics are often unavailable and hard to aggregate. Preference learning provides an alternative by learning from comparative human or AI feedback. Yet, its extension to multi-agent systems remains underexplored. To address this gap, we formulate preference-based multi-agent systems (MAS) from decentralized and centralized collaboration perspectives. We also introduce a general multi-agent preference learning framework (MAPL) to solve these problems. MAPL allows iterative updates by comparing the current solution with decentralized or centralized solutions generated by various agents. We instantiate MAPL using MARL from human feedback (MARLHF) with a learned reward model and multi-agent direct preference optimization (MADPO). Experiments on collaborative writing, coding, tool use, and travel planning show that MAPL can improve collaboration quality and efficiency while approaching the performance of MARL with fixed, well-defined rewards. Within MAPL, MARLHF generally outperforms MADPO on most tasks but remains sensitive to data coverage, agent and comparator models, and the underlying MARL algorithms.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Systems
Preference Learning
Large Language Models
Collaboration
Multi-Agent Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Preference Learning
MAPL
Multi-Agent Reinforcement Learning from Human Feedback (MARLHF)
Multi-Agent Direct Preference Optimization (MADPO)
LLM Collaboration
🔎 Similar Papers