🤖 AI Summary
This study addresses the challenges of adaptation and coordination faced by multi-agent offline flow policies in unseen environments. To overcome these limitations, this work proposes a fine-tuning framework that integrates flow matching pretraining with online reinforcement learning. Specifically, the method constructs an explicit action-likelihood policy and introduces an update mechanism based on shared team advantages, enabling efficient optimization of cooperative behaviors across both discrete and continuous action spaces. This approach effectively resolves the difficulties of cross-environment generalization and multi-agent coordination. Empirical evaluations demonstrate substantial improvements, achieving a 52.8% average performance gain over the strongest offline baseline and a 29.8% improvement compared to purely online learning methods.
📝 Abstract
Multi-agent flow policies learn cooperative behavior from fixed offline datasets, but often struggle to complete tasks in situations not covered by the offline data. In these situations, agents must both adapt to changes in the environment and coordinate with one another, yet action patterns learned offline are often insufficient for effective adaptation and coordination. To address this problem, we propose Multi-Agent Flow-Pretrained Policy Optimization (MA-FPPO), which uses online fine-tuning to improve the cooperative behavior of models pretrained with flow matching through new interactions with the environment. Building on the behavior learned during pretraining, we construct policies with explicit action likelihoods for discrete and continuous action spaces. We then update the pretrained model using shared team advantages to further improve coordination based on team performance. Our method achieves, on average, relative gains of 52.8% over the strongest listed offline baselines across 30 settings and 29.8% over purely online learning across 38 comparisons with matched online budgets and evaluation protocols.