🤖 AI Summary
This study addresses the high cost and deployment complexity associated with force-aware manipulation relying on dedicated sensors by proposing a vision-based force estimation method that eliminates the need for tactile or force sensors. Specifically, contact forces are predicted by observing the deformation of a compliant gripper. Furthermore, an action-force joint policy network is constructed to generate candidate actions alongside their expected force feedback, enabling optimal action selection during inference through target-force-guided sampling. This approach facilitates contact-rich robotic manipulation without requiring additional sensing hardware. The effectiveness of the proposed method is validated across practical tasks, including berry harvesting and aluminum can grasping, demonstrating its potential for accessible and robust force-sensitive manipulation in unstructured environments.
📝 Abstract
Force-aware manipulation typically relies on specialized force or tactile sensors. We show that force-aware manipulation can instead be achieved through visual force prediction from the deformation of a compliant Fin Ray gripper. Our approach trains two models. First, we train a visual force estimator on calibration data and use it to annotate task demonstrations with force estimates. Second, we train an action--force proposal policy on these force-augmented demonstrations to jointly generate candidate robot actions and their associated forces. At test time, we sample candidate actions and the forces they are expected to produce, then execute the action whose predicted force is closest to a target from the demonstrations. We evaluate our approach on berry picking, empty-can grasping, in-hand reorientation, and plug insertion. Our results show that visual force prediction can guide inference-time action selection for contact-rich manipulation without requiring force or tactile sensors at deployment.