An Empirical Study and Open Testbed for Federated Fine-Tuning of Vision-Language-Action Models

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文研究联邦微调方法解决预训练视觉-语言-动作模型在新机器人、环境或任务中的适应问题,通过实证分析多种参数和算法选择的影响。
📝 Abstract
Adapting a pretrained Vision-Language-Action (VLA) model to a new robot, environment, or task requires demonstrations that are collected locally and often discarded. Federated learning is a promising approach to exploiting such distributed demonstrations by learning a shared policy. However, whether it can adapt large pretrained VLAs remains an open question, and a lack of reproducible benchmarks for pretrained VLAs and reusable training frameworks makes existing results difficult to compare. In this paper, we conduct a systematic study of federated fine-tuning of three modern pretrained VLA policies on the 40 simulated tasks of the LIBERO manipulation benchmark, and on six real-world tasks in two real-robot experiments, with demonstrations collected across two and three sites, respectively. Our study analyzes the key choices in this setting, spanning multiple federated parameter scopes, three aggregation algorithms, and evaluation under distribution shift. Based on the study, we derive a series of lessons, including the dominance of the federated scope over the choice of aggregation algorithm and the difficulty of matching centralized fine-tuning on physical robots, where cross-site heterogeneity is stronger than simulation captures. We also highlight opportunities for federated VLA learning, such as the ability to match centralized fine-tuning on heterogeneous data, to remain at least as robust as centralized fine-tuning under distribution shift, and to personalize, with each client federating part of the policy and keeping the rest local, which helps where the policy's pretraining is weak but leaves no usable global model. We open-source \decentvla{}, the model- and runtime-agnostic testbed behind the study, to facilitate future research and fair comparisons in federated VLA learning.
Problem

Research questions and friction points this paper is trying to address.

Federated Learning
Vision-Language-Action Models
Distributed Demonstrations
Adaptation
Benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated Fine-Tuning
Vision-Language-Action Models
Distributed Demonstrations
Aggregation Algorithms
Personalization
🔎 Similar Papers
Z
Zhekai Duan
Department of Computer Science, University College London, U.K.
K
Kevin Ziyang Xie
Department of Computer Science, University College London, U.K.
X
Xinyu Tan
Department of Computer Science, University College London, U.K.
S
Shikai Geng
Department of Computer Science, University College London, U.K.
Chengxu Zhou
Chengxu Zhou
Associate Professor in Robotics & AI, University College London
Legged ManipulationWhole Body ControlHumanoid RobotTelexistenceEmbodied AI
R
Ramana Kompella
Cisco Research, San Jose, CA, USA.
Gaowen Liu
Gaowen Liu
Cisco Research
machine learningcomputer visionmultimedia.
Chris Xiaoxuan Lu
Chris Xiaoxuan Lu
Associate Professor at University College London (UCL)
RoboticsCyber Physical Systems