🤖 AI Summary
This study addresses the high sampling costs and limited online scalability of diffusion policies in multi-agent reinforcement learning by proposing the OMAF framework. The method introduces an approximate path score surrogate and designs a Transformer-based flow model equipped with a one-step generation mechanism. Furthermore, a joint optimization scheme couples Softmax Q-value estimation with the flow objective, effectively eliminating iterative sampling overhead while enabling efficient and expressive coordination policy learning. Evaluated on benchmark tasks such as MPE, OMAF achieves a 3.4-fold improvement in returns and a 10.5-fold increase in sample efficiency, significantly outperforming existing baseline methods.
📝 Abstract
Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies, particularly diffusionbased policies, can faithfully capture complex and multimodal behaviors, but costly iterative sampling hinders their scalability in online multi-agent settings. We propose an Online MARL framework via one-step Flow model (OMAF) that combines expressive generative policies with efficient one-step action generation. OMAF employs a Transformer-based flow policy to capture complex coordination behaviors, while its approximate path score surrogate provides a principled route to synchronized flow policy optimization. To enable stable and sampleefficient learning, we further develop a joint optimization scheme coupling softmax Q-value estimation with a joint flow policy objective for coordinated policy learning. By eliminating iterative sampling, OMAF dramatically reduces training overhead without sacrificing policy expressiveness. Extensive experiments across 10 standard tasks from MPE and MAMuJoCo show that OMAF consistently achieves superior performance, with up to 3.4x higher returns and 10.5x sample efficiency improvement compared with baseline methods. These results validate the effectiveness of OMAF as an expressive and computationally efficient one-step flow policy paradigm for online MARL.