Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited collaborative capabilities and constrained supervised fine-tuning of multi-agent Vision-Language-Action (VLA) models by proposing a three-stage reinforcement fine-tuning framework. The method pioneers latent-space online reinforcement learning tailored for VLA models, effectively circumventing noisy exploration. Furthermore, it optimizes offline tuning through initialization-aware data collection and credit assignment mechanisms, enabling efficient training from single-agent trajectories. Evaluated on the RoboTwin benchmark and real-world dual-Franka robot tasks, the proposed approach increases average success rates by 23.1%, 16.4%, and 44%, respectively, substantially enhancing multi-robot cooperative control performance.
📝 Abstract
We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to initialization shift with reduced human cost. Second, offline credit-filtered tuning assigns credit to individual agents and fine-tunes on per-agent trajectories with positive advantage rather than on entire joint rollouts. Third, we find existing online RL for VLAs are less effective for hard multi-agent tasks, which we attribute to noisy co-exploration and unstable updates. We instead use online latent-space fine tuning, which freeze the VLA and perform RL in its latent noise space. We evaluate our multi-agent VLA with both $π_0$ and $π_{0.5}$ backbones across 11 tasks in RoboTwin, RoboFactory and real-world manipulation with two Franka robots. Our multi-agent VLA improves the average success rate by $+23.1\%$, $+16.4\%$, and $+44\%$ on RoboTwin, RoboFactory, and real-world tasks, respectively. Code available at https://anonymous.4open.science/r/mavla_rft-2BC0/.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent VLA
Cooperative Robots
Reinforcement Learning
Vision-Language-Action Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent VLA
Reinforced Fine-Tuning
Credit Assignment
Latent-Space RL
Cooperative Manipulation
🔎 Similar Papers
R
Ruixiao Xu
Beihang University
W
WONG Lik Hang Kenny
The Chinese University of Hong Kong
Z
Zhiqian Liu
Beihang University
J
Jianing Guo
Beihang University, PKU-Psibot Lab
H
Hanxiao Li
Beihang University
K
Kejian Shi
The Chinese University of Hong Kong
Shuning Zhang
Shuning Zhang
Tsinghua University
HCIUsable Privacy and SecurityAI
P
Pu Feng
Beihang University, Zhongguancun Laboratory
Yongjia Ma
Yongjia Ma
LiAuto
Computer VisionNeural RenderAIGCVLA
Y
Yuqing Ma
Beihang University
Kai Chen
Kai Chen
Hong Kong University of Science and Technology
Representation LearningGenerative ModelingMulti-modalityMixture-of-Experts
Q
Qi Dou
The Chinese University of Hong Kong
Yaodong Yang
Yaodong Yang
Boya (博雅) Assistant Professor at Peking University
Reinforcement LearningAI AlignmentEmbodied AI
X
Xianglong Liu
Beihang University, Zhongguancun Laboratory
Simin Li
Simin Li
Beihang University
Reinforcement LearningMulti-Agent LearningAdversarial attackTrustworthy AI