🤖 AI Summary
Existing robot-to-human object handover methods neglect post-handover human manipulation requirements, rely on strong assumptions, and exhibit poor generalizability. This paper proposes the first semantic-driven handover framework integrating large language models (LLMs) with vision-based part perception: an LLM interprets task intent and infers functionally critical object parts, while RGB-D–based instance segmentation and grasp pose estimation enable context-aware grasp selection. To support zero-shot grasp planning, we introduce a fine-grained part-annotated dataset covering 60 household object categories. Real-robot experiments achieve an 83% grasp success rate; in user studies, 86% of participants preferred our method, demonstrating significant improvements over baselines in both naturalness and task adaptability.
📝 Abstract
Effective human-robot collaboration depends on task-oriented handovers, where robots present objects in ways that support the partners intended use. However, many existing approaches neglect the humans post-handover action, relying on assumptions that limit generalizability. To address this gap, we propose LLM-Handover, a novel framework that integrates large language model (LLM)-based reasoning with part segmentation to enable context-aware grasp selection and execution. Given an RGB-D image and a task description, our system infers relevant object parts and selects grasps that optimize post-handover usability. To support evaluation, we introduce a new dataset of 60 household objects spanning 12 categories, each annotated with detailed part labels. We first demonstrate that our approach improves the performance of the used state-of-the-art part segmentation method, in the context of robot-human handovers. Next, we show that LLM-Handover achieves higher grasp success rates and adapts better to post-handover task constraints. During hardware experiments, we achieve a success rate of 83% in a zero-shot setting over conventional and unconventional post-handover tasks. Finally, our user study underlines that our method enables more intuitive, context-aware handovers, with participants preferring it in 86% of cases.