🤖 AI Summary
This study addresses the issue that grasping diffusion models often ignore embodied constraints, yielding physically infeasible generations. To this end, we propose a training-free, inference-time guidance framework. Specifically, it introduces a novel Feynman-Kac-inspired particle resampling mechanism, integrated with reachability- and collision-aware rewards alongside gradient-free local refinement. This approach effectively shifts probability mass from infeasible modes toward high-reward feasible regions, achieving population-level mode transitions without modifying the frozen pretrained model or requiring differentiable constraints. Experimental results demonstrate that the proposed method substantially improves the generation rate of physically feasible grasps across diverse object and robot scenarios while maintaining strong fidelity to the original prior distribution.
📝 Abstract
Current grasp diffusion models provide rich priors for generation, yet their object-centric approach can violate the kinematic and collision constraints imposed by the embodiment and the environment. Existing embodiment-aware methods primarily perform local corrections around generated grasps through gradient guidance or optimization, making it difficult to recover from fundamentally infeasible modes. We present Steer2Grasp, a training-free, embodiment-agnostic framework for inference-time grasp steering that adapts a frozen Cartesian grasp diffusion model using deployment-specific rewards. Through Feynman-Kac (FK) inspired particle reweighting and resampling, the method reallocates population mass from infeasible to high-reward grasp modes, enabling population-level mode transitions without modifying the pretrained diffusion model or requiring differentiable constraints. The framework enables a unified treatment for single and dual arm grasping through reachability and collision aware rewards, followed by gradient free gripper level local refinement. Across diverse objects, robot embodiments, and constrained environments, our method substantially improves feasible grasp generation while maintaining proximity to the underlying grasp prior.