Co-VLA: Consensus-based Federated Training for Vision-Language-Action Models

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Co-VLA,使用ADMM共识优化方法解决联邦学习中视觉-语言-动作模型训练的数据异构问题,实现无需数据共享的协作训练。
📝 Abstract
Vision-language-action models (VLAs) have emerged as a promising paradigm for general-purpose robot learning, with performance improving as models and datasets scale. Scaling robot data collection, however, remains challenging because data are naturally distributed across robots, tasks, and locations, making centralization costly or impractical. Federated learning offers a way to train on decentralized robot data, but applying it to VLAs requires accounting for heterogeneous robot client data distributions. We present Co-VLA, which applies consensus optimization using the Alternating Direction Method of Multipliers~(ADMM) to federated VLA training. We show that the same algorithm supports both full-model training and parameter-efficient fine-tuning with both fixed-rank and rank-adaptive adapters. The name Co-VLA reflects both consensus and collaboration: clients with different local robot datasets collaboratively train a shared model without sharing their data. Our experiments demonstrate that Co-VLA achieves performance comparable to centralized training in both full-model training and parameter-efficient fine-tuning settings.
Problem

Research questions and friction points this paper is trying to address.

Federated Learning
Vision-Language-Action Models
Heterogeneous Data Distributions
Robot Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Consensus Optimization
Federated Learning
Vision-Language-Action Models
ADMM
Parameter-efficient Fine-tuning
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
H
Haolong Li
Intelligent Perception in Technical Systems, University of Augsburg, Germany
G
Guner Dilsad Er
Max Planck Institute for Intelligent Systems, Germany
Michael Muehlebach
Michael Muehlebach
Max Planck Institute for Intelligent Systems
Machine LearningOptimizationDynamical Systems
J
Joerg Stueckler
Intelligent Perception in Technical Systems, University of Augsburg, Germany