🤖 AI Summary
This study addresses the challenge of collecting demonstration data for long-horizon dexterous manipulation that simultaneously captures macro-level task progression and micro-level contact interactions. To this end, it proposes a novel heterogeneous demonstration framework that matches interaction modalities by integrating teleoperation with kinesthetic teaching. Furthermore, a mask-conditioned diffusion policy supervised via offline segmentation is designed to resolve visual mismatches, while a successor-aware steering algorithm is introduced to enable smooth policy transitions and mitigate distribution shift. Experimental results demonstrate an end-to-end success rate of 27%, with dexterous subtask success rates improving to 65% and the average policy composition efficiency reaching 87%.
📝 Abstract
Dexterous manipulation requires both large-scale task progression and precise contact-rich interaction, making it challenging to collect demonstrations that effectively support both regimes. We present SkillWeave, a heterogeneous demonstration framework for long-horizon dexterous manipulation that combines teleoperation for coarse reaching and transport with kinesthetic teaching for precise, contact-rich skills. To address the visual mismatch introduced by the demonstrator's presence during kinesthetic data collection, we propose an object-mask-conditioned diffusion policy that uses offline object segmentation for training supervision and a lightweight learned mask predictor at deployment, avoiding online segmentation and image inpainting. To mitigate distribution shift between independently trained sub-task policies, we introduce successor-aware terminal steering, which selects among actions sampled from the predecessor policy to guide the system toward states supported by the successor's demonstrated initial-state distribution. Across three real-world long-horizon tasks, SkillWeave achieves 27% average end-to-end success. Mask-conditioned kinesthetic policies improve dexterous sub-task success to an average of 65%, while successor-aware handoffs achieve an average composition efficiency of 87%. These results show that matching demonstration modality to interaction regime, explicitly addressing kinesthetic visual mismatch, and steering policy handoffs toward successor-supported states substantially improves long-horizon dexterous manipulation. Videos and code are available at skillweave-authors.github.io .