Score
Designs and implements pretraining and transfer-learning pipelines, curricula, datasets, and evaluation methods to enable models and control policies to generalize across different embodiments (for example human and robot bodies or hands), including mixed human–robot pretraining, progressive human-to-robot adaptation, multi-embodiment and multi-robot pretraining, and zero-shot transfer. This work builds mixed-data corpora (videos, demonstrations, trajectories), constructs progressive/domain-adaptation schedules and co-training schemes, and develops embodiment-agnostic policy-learning and cross-embodiment evaluation procedures.
This work addresses the challenge of data scarcity in robot learning by providing a systematic survey of recent advances in transferring manipulation skills from human videos. It introduces the first hierarchical taxonomy tailored to robotic skill acquisition, integrating human-to-robot transfer pathways, data configurations, and learning paradigms across three levels: task, observation, and action. The study further presents a large-scale statistical analysis of existing video datasets, characterizing their scale, structure, and usage trends. By synthesizing developments in policy learning, computer vision, generative modeling, and cross-paradigm coupling methods, this survey comprehensively maps the current landscape, identifies key challenges and limitations, and outlines promising future directions. To foster community progress, the authors also release an open-source collection of relevant papers.
Current autoregressive vision-language-action (VLA) models exhibit limited generalization across heterogeneous robotic platforms in multi-robot collaborative tasks. To address this, we propose ET-VLA, an entity-transfer learning framework that integrates continual pretraining on synthetic data with target-entity fine-tuning, enabling efficient cross-morphology and cross-robot adaptation. We further introduce an embodied reasoning graph to explicitly model functional roles of multiple entities and their interdependent action relationships, supporting fine-grained collaborative action generation. Our approach significantly reduces reliance on scarce real-world demonstration data while enhancing adaptability to diverse robot systems. Evaluated on three real-world dual-arm robotic tasks, ET-VLA achieves an average performance gain of 53.2% over OpenVLA, demonstrating strong sim-to-real generalization. The code is publicly available.
Robot manipulation policies exhibit poor generalization across morphologically distinct embodiments and lack standardized evaluation protocols. Method: We propose the first benchmark for cross-morphology manipulation tasks, covering fundamental grasping and pushing tasks, and enabling systematic assessment of interpolation, extrapolation, and composition-based generalization. We introduce a three-axis evaluation framework that formally defines and quantifies “morphology-agnostic manipulation policy generalization.” Our approach employs a morphology-aware policy architecture trained via multi-morphology joint reinforcement learning and grounded in structured simulation modeling to enable zero-shot transfer. Contribution/Results: This benchmark fills a critical gap in the field. Experiments reveal substantial performance degradation under morphological extrapolation; morphology-aware training consistently outperforms single-morphology baselines, yet zero-shot generalization across structurally divergent embodiments remains a fundamental challenge.
This study investigates how to effectively organize heterogeneous robotic demonstration data to enhance cross-embodiment transfer performance. Through controlled simulation experiments, it systematically compares the efficacy of unpaired large-scale data against structured paired data—such as demonstrations aligned by scene, task, or trajectory—under varying morphologies and viewpoints. The findings reveal that, for morphology differences, structured data analogies are more effective than merely increasing data diversity, highlighting distinct data structure requirements for morphology transfer versus viewpoint transfer. By optimizing data composition alone, the approach achieves an average 22.5% improvement in success rate on real-world cross-embodiment transfer tasks, underscoring the critical role of data analogy in enabling effective embodiment-agnostic skill transfer.
This work addresses the limited generalization of existing vision-language-action models across diverse robot morphologies and data-scarce scenarios. The authors propose a human-centric cross-embodiment learning paradigm that treats human interaction actions as a universal “lingua franca,” enabling seamless transfer from human demonstrations to robot execution. By mapping heterogeneous robot controls into a unified action space grounded in human motion, the approach integrates sequence modeling with multi-task pretraining. Key innovations include a Mixture-of-Flow architecture, a manifold-preserving gating mechanism, and a universal asynchronous chunking strategy. Evaluated on the UniHand-2.0 multimodal pretraining dataset, the model achieves state-of-the-art performance on the LIBERO (98.9%) and RoboCasa (53.9%) simulation benchmarks and demonstrates exceptional cross-embodiment generalization across five real-world robotic platforms.
This work addresses the challenge of policy transfer across robotic morphologies, which is hindered by embodiment differences that limit the generalization of imitation learning. To overcome this, the authors propose leveraging embodiment-invariant behavioral alignment representations—such as end-effector trajectories, object bounding boxes, and language-based action descriptions—to construct a unified vision-language-action (VLA) model capable of integrating multi-embodiment data and enabling effective cross-embodiment transfer. The approach presents the first systematic evaluation of behavioral alignment representations in sim-to-real transfer, demonstrating substantial performance gains on a newly introduced simulation benchmark. When deployed on real robots, the method improves task success rates by 28% and further enables training augmentation using unlabeled demonstration data lacking action annotations.
Cross-configurational robot learning suffers from limited reusability of data, policies, and control code due to highly customized and fragmented infrastructure. To address this challenge, this work proposes RIO—an open-source, lightweight, and modular Python framework that enables flexible switching across hardware platforms and tasks through a unified abstraction layer. RIO supports multi-platform robot control, teleoperation, sensor configuration, data standardization, and deployment of vision–language–action (VLA) policies. The framework has been validated on three distinct robot morphologies and four hardware platforms, significantly lowering the barrier to cross-platform reuse. It has also been successfully employed to fine-tune state-of-the-art VLA models such as π₀.₅ and GR00T, enabling them to perform diverse household tasks including grasping, folding, and dishwashing, thereby advancing the ecosystem for general-purpose robot learning.
This work addresses the limited generalizability of existing cross-embodiment video generation methods, which suffer from entangled motion and morphology representations and rely on paired data for target embodiments. To overcome these limitations, the authors propose a motion-morphology disentangled modeling framework that enables rapid adaptation to new robots without requiring paired data, leveraging a shared motion model and lightweight embodiment adapters. A novel branch-isolated attention mechanism is introduced to effectively separate motion conditioning from embodiment-specific modulation. The study also presents the first large-scale synthetic dataset of cross-embodiment paired videos. Experimental results demonstrate high motion fidelity and embodiment consistency on both synthetic and real-world benchmarks, with successful zero-shot transfer to unseen humanoid embodiments without retraining the shared motion model.