Beyond the Current Scene: Event-Referential Grasping with Active View Selection

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of enabling robots to grasp potentially occluded target objects based on historical event references. To this end, it proposes BeyondSCe, a zero-shot grasping system that eliminates the need for task-specific training. The approach integrates event memory reasoning with scene geometric priors, leveraging pretrained models and RGB-D vision to achieve precise localization and manipulation of occluded targets via an active viewpoint selection algorithm. A key contribution lies in the effective fusion of event priors with active perception. Experimental results demonstrate that the system attains grasping success rates of 76% and 77% for visible and occluded objects, respectively. Notably, in heavily occluded scenarios, the success rate improves to 95%, accompanied by a significant reduction in the number of required viewpoint switches.
📝 Abstract
A robot that observes people interacting with objects should be able to carry out later requests that refer back to those interactions. Such requests may specify a grasp target by the role it played in a past event rather than by its name or appearance. Moreover, the target may no longer be visible when the robot is asked to act. We present BeyondSCe, a zero-shot robotic grasping system for this event-referential setting. Given the event history and the current scene, the system identifies the requested object or part and localizes it for grasping. If the target is occluded, it combines an event prior recovered from the history with current scene geometry to select camera viewpoints likely to reveal the target. The system uses pretrained models without additional task-specific training. In real-robot experiments with a single wrist-mounted RGB-D camera, it achieves grasp success rates of 76% and 77% for initially visible and occluded targets, respectively, compared with 40% and 55% for the strongest baseline in each condition. On four additional scenes with heavy occlusion, it increases grasp success rates from 75% to 95% while reducing the mean number of views from 3.35 to 2.20, compared with an active-perception baseline given the target's ground-truth 3D bounding box.
Problem

Research questions and friction points this paper is trying to address.

event-referential grasping
active view selection
robotic grasping
occlusion
zero-shot
Innovation

Methods, ideas, or system contributions that make the work stand out.

Event-Referential Grasping
Zero-Shot Robotic Grasping
Active View Selection
Occlusion Reasoning
Pretrained Models