DexJoCo-X: Benchmarking Action Representations for Multi-Hand Dexterous Manipulation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of cross-embodiment learning for dexterous robotic hands, where hardware, task, and interface discrepancies hinder the disentanglement of action representation effects. To overcome this, we construct a unified benchmark encompassing seven dexterous hand morphologies across multiple tasks and propose a standardized evaluation protocol. By aligning scene execution interfaces, redesigning glove-based mappings, and developing an automated data augmentation pipeline, we systematically assess the synergistic interactions among action representations, pretraining strategies, and network architectures. Our findings reveal that shared structures necessitate holistic co-design. Notably, we demonstrate that functionally aligned action slots achieve a 47.7% average success rate, substantially outperforming native coordinate and latent representation baselines. These results offer critical insights into designing generalizable manipulation policies across diverse dexterous embodiments.
📝 Abstract
As dexterous hands proliferate, collecting data and training policies separately for every morphology becomes increasingly impractical. Scalable cross-embodiment learning therefore requires a unified representation that captures shared manipulation structure while preserving morphology-specific control. Differences in hands, tasks, datasets, and control interfaces prevent existing studies from isolating the effects of representation, pretraining, and architecture. We introduce DexJoCo-X, a benchmark and toolkit for controlled comparison across seven representative dexterous hands, six single-arm and bimanual tasks, and 2,100 balanced demonstrations. DexJoCo-X provides a matched multi-hand, multi-task protocol with common scenes, success criteria, and execution interfaces, redesigned glove-to-hand mappings, and an automated pipeline that expands reviewed demonstrations across randomized scenes. Using $π_{0.5}$, Ego-Pi, and Being-H0.5, we examine whether a shared action interface is sufficient for multi-hand learning. Expanding $π_{0.5}$ to an 80-dimensional bimanual output yields near-zero success. Ego-Pi preserves the pretrained action head through interleaved prediction and supports per-hand multi-task learning, but remains ineffective for seven-hand joint training. By contrast, Being-H0.5 combines cross-embodiment pretraining, a unified action space, and embodiment-aware experts, enabling one policy to control all seven hands. Within this architecture, function-aligned action slots achieve 47.7% mean success, compared with 47.0% for native coordinates and 33.1% for DexLatent. These results show that cross-embodiment representation depends on the entire learning system: action coordinates, pretraining, and architecture must jointly separate shared manipulation structure from embodiment-specific control.
Problem

Research questions and friction points this paper is trying to address.

dexterous manipulation
cross-embodiment learning
action representation
multi-hand benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-embodiment learning
dexterous manipulation
action representation
benchmark
embodiment-aware experts
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiangwei Jiang
University of Electronic Science and Technology of China, China
Y
Yao Mu
Shanghai Jiao Tong University, China
Lixin Duan
Lixin Duan
Data Intelligence Group (DIG) @ UESTC
Transfer LearningDomain Adaptation
W
Wen Li
University of Electronic Science and Technology of China, China