Embodiment Transfer Learning for Vision-Language-Action Models

📅 2025-11-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current autoregressive vision-language-action (VLA) models exhibit limited generalization across heterogeneous robotic platforms in multi-robot collaborative tasks. To address this, we propose ET-VLA, an entity-transfer learning framework that integrates continual pretraining on synthetic data with target-entity fine-tuning, enabling efficient cross-morphology and cross-robot adaptation. We further introduce an embodied reasoning graph to explicitly model functional roles of multiple entities and their interdependent action relationships, supporting fine-grained collaborative action generation. Our approach significantly reduces reliance on scarce real-world demonstration data while enhancing adaptability to diverse robot systems. Evaluated on three real-world dual-arm robotic tasks, ET-VLA achieves an average performance gain of 53.2% over OpenVLA, demonstrating strong sim-to-real generalization. The code is publicly available.

Technology Category

Intelligent Robots: Embodied AIComputer Vision: Multi-modal VisionHumans and AI: Human-Aware Planning and Behavior Prediction

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomySemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
📝 Abstract
Vision-language-action (VLA) models have significantly advanced robotic learning, enabling training on large-scale, cross-embodiment data and fine-tuning for specific robots. However, state-of-the-art autoregressive VLAs struggle with multi-robot collaboration. We introduce embodiment transfer learning, denoted as ET-VLA, a novel framework for efficient and effective transfer of pre-trained VLAs to multi-robot. ET-VLA's core is Synthetic Continued Pretraining (SCP), which uses synthetically generated data to warm up the model for the new embodiment, bypassing the need for real human demonstrations and reducing data collection costs. SCP enables the model to learn correct actions and precise action token numbers. Following SCP, the model is fine-tuned on target embodiment data. To further enhance the model performance on multi-embodiment, we present the Embodied Graph-of-Thought technique, a novel approach that formulates each sub-task as a node, that allows the VLA model to distinguish the functionalities and roles of each embodiment during task execution. Our work considers bimanual robots, a simple version of multi-robot to verify our approaches. We validate the effectiveness of our method on both simulation benchmarks and real robots covering three different bimanual embodiments. In particular, our proposed ET-VLA space can outperform OpenVLA on six real-world tasks over 53.2%. We will open-source all codes to support the community in advancing VLA models for robot learning.
Problem

Research questions and friction points this paper is trying to address.

Enables efficient transfer of vision-language-action models to multi-robot systems
Reduces data collection costs through synthetic pretraining without human demonstrations
Enhances multi-robot collaboration by distinguishing embodiment roles during tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Embodiment transfer learning for multi-robot collaboration
Synthetic Continued Pretraining bypasses human demonstration needs
Embodied Graph-of-Thought distinguishes embodiment roles in tasks
🔎 Similar Papers
2024-05-14IEEE/RJS International Conference on Intelligent RObots and SystemsCitations: 2
💼 Related Jobs
No related jobs found.
C
Chengmeng Li
Shanghai University
Y
Yaxin Peng
Shanghai University