π€ AI Summary
This study addresses the limited robustness of multimodal object handovers caused by robots' inability to anticipate task-specific hand poses. To this end, we propose GENESIS-Handover, a framework that pioneers the integration of vision-language model (VLM) image generation into handover systems. By employing generative hypothesis selection, our method explicitly models the task-specific variability of hand poses. Specifically, it leverages VLMs to generate hand-object interaction hypotheses and matches them against observed poses in real time, thereby inferring task-oriented, intelligent handover strategies. A user study demonstrates that 83.3% of participants perceived the proposed approach as significantly outperforming existing state-of-the-art methods in task understanding capabilities.
π Abstract
When humans hand each other objects, they incorporate both geometric and semantic information into this process. For example, passing a knife with the handle towards the recipient, rather than the blade, is both more ergonomic and safer. Recent state-of-the-art methods for task-oriented robot-human handovers have progressed from modeling object geometry to incorporating object affordances. However, they often forgo predicting the explicit, task-specific hand poses a human selects to utilize an object. Since many objects support multiple interaction modalities, e.g., a claw hammer used to strike or pull nails, this variability must be modeled to achieve robust task-oriented handovers. To tackle this, we propose a novel approach, GENESIS-Handover (GENErative HypotheSIS), which leverages VLM image generation to produce a variety of task-specific hand-object interaction hypotheses. These hypotheses are matched in real time to the observed human hand pose, enabling inference of the most suitable handover configuration. By leveraging VLMs as priors of plausible hand-object interactions, the method produces task-conditioned handover strategies for previously unseen object-task pairs. We evaluate the standalone interaction proposal module before deploying the full system on a mobile manipulator. In a user study with 12 participants across five task-object pairs, 83.3% perceived our method to have better task understanding than the previous state of the art.