🤖 AI Summary
This study addresses the substantial retraining requirements and poor cross-domain generalization associated with robotic grasping in novel scenarios by proposing a retrieval-augmented, retraining-free adaptation framework for parallel-jaw grippers. Methodologically, a two-level cascaded geometric-semantic retrieval mechanism with confidence gating is designed to enable few-shot scene adaptation. Furthermore, 2D-to-3D grasp transfer is achieved by integrating DINOv2 features, SAM2 segmentation, and mask-based constraints. Real-world experiments demonstrate that the proposed framework attains grasping success rates of 100% and 95% on known and unknown objects, respectively, while exhibiting strong robustness against viewpoint and illumination variations.
📝 Abstract
We present RAGrasp, a retrieval-augmented pipeline for planar parallel-jaw grasping from a compact set of locally collected, grasp-annotated RGB-D (color and depth) templates. Unlike task-specific predictors trained primarily on large public or synthetic grasp datasets, RAGrasp requires no end-to-end retraining for a new deployment.Its template memory is constructed from observations collected with the deployment camera, robot, and gripper in the target workspace, thereby aligning stored examples with the local sensing and embodiment conditions. The system uses self-supervised DINOv2 visual fea- tures together with appearance and depth cues to prompt the Segment Anything Model 2 (SAM2), which isolates the query object. A two-stage geometry-semantic retrieval cascade then selects a template, and a confidence gate chooses one of two grasp- transfer estimators. The transferred grasp is refined using mask- support and silhouette-contact constraints before calibrated 2D- to-3D conversion. In real-world trials, RAGrasp achieves 20/20 successful grasps on seen objects and 19/20 on unseen objects. Within the evaluated setting, the results demonstrate deployment- specific grasp adaptation from limited local annotation and tolerance to the tested viewpoint and illumination changes.