🤖 AI Summary
This study addresses the challenge that dexterous hands lack reusable, cross-task and cross-embodiment contact primitives, making it difficult to translate manipulation intent into stable multi-finger behaviors. To this end, we propose the GOTT framework, which follows an “approach–acquire–transport” paradigm. This method introduces the first unified, cross-embodiment closed-loop contact acquisition primitive for establishing stable grasps, and integrates a pose-conditioned controller with robot-agnostic trajectory planning to achieve robust manipulation. Its core contribution lies in decoupling high-level planning from low-level control, thereby supporting multi-source approach specifications without requiring backend modifications. Simulation and real-world experiments demonstrate that GOTT reliably establishes robust contacts across diverse objects, different arm-hand platforms, and unseen morphologies, significantly improving end-to-end task success rates.
📝 Abstract
Foundation models and large-scale human data provide rich sources of manipulation intent, but translating this intent into multi-fingered robot behavior remains difficult. Dexterous hands still lack a reusable low-level primitive that reliably establishes contact across tasks and embodiments. We propose GOTT, a reach-acquire-move framework built around a single cross-embodiment contact-acquisition primitive. Given a robot-agnostic object trajectory and a reach specification, GOTT first brings the hand near a task-relevant contact region. The shared closed-loop primitive then establishes stable contact from this approximate initialization, and a pose-conditioned controller tracks the desired object motion. Reach specifications may come from future-aware planning, external models, or human demonstrations, while the primitive and tracking backend remain unchanged. Simulation and real-world experiments show that GOTT is able to establish robust contact across diverse objects, arm-hand platforms, and seen and unseen hand morphologies. It also consistently improves end-to-end task success over open-loop grasp execution.