🤖 AI Summary
This study addresses the challenge of cross-embodiment learning for dexterous robotic hands, where hardware, task, and interface discrepancies hinder the disentanglement of action representation effects. To overcome this, we construct a unified benchmark encompassing seven dexterous hand morphologies across multiple tasks and propose a standardized evaluation protocol. By aligning scene execution interfaces, redesigning glove-based mappings, and developing an automated data augmentation pipeline, we systematically assess the synergistic interactions among action representations, pretraining strategies, and network architectures. Our findings reveal that shared structures necessitate holistic co-design. Notably, we demonstrate that functionally aligned action slots achieve a 47.7% average success rate, substantially outperforming native coordinate and latent representation baselines. These results offer critical insights into designing generalizable manipulation policies across diverse dexterous embodiments.
📝 Abstract
As dexterous hands proliferate, collecting data and training policies separately for every morphology becomes increasingly impractical. Scalable cross-embodiment learning therefore requires a unified representation that captures shared manipulation structure while preserving morphology-specific control. Differences in hands, tasks, datasets, and control interfaces prevent existing studies from isolating the effects of representation, pretraining, and architecture. We introduce DexJoCo-X, a benchmark and toolkit for controlled comparison across seven representative dexterous hands, six single-arm and bimanual tasks, and 2,100 balanced demonstrations. DexJoCo-X provides a matched multi-hand, multi-task protocol with common scenes, success criteria, and execution interfaces, redesigned glove-to-hand mappings, and an automated pipeline that expands reviewed demonstrations across randomized scenes. Using $π_{0.5}$, Ego-Pi, and Being-H0.5, we examine whether a shared action interface is sufficient for multi-hand learning. Expanding $π_{0.5}$ to an 80-dimensional bimanual output yields near-zero success. Ego-Pi preserves the pretrained action head through interleaved prediction and supports per-hand multi-task learning, but remains ineffective for seven-hand joint training. By contrast, Being-H0.5 combines cross-embodiment pretraining, a unified action space, and embodiment-aware experts, enabling one policy to control all seven hands. Within this architecture, function-aligned action slots achieve 47.7% mean success, compared with 47.0% for native coordinates and 33.1% for DexLatent. These results show that cross-embodiment representation depends on the entire learning system: action coordinates, pretraining, and architecture must jointly separate shared manipulation structure from embodiment-specific control.