🤖 AI Summary
This study addresses the reliance of contact-rich manipulation tasks on task-specific training or handcrafted rules by proposing the first zero-shot vision-language model (VLM) and tactile fusion framework for training-free autonomous peg insertion. The proposed method integrates a pretrained VLM with multimodal inputs, comprising visual observations and tri-axial tactile signals, to directly generate action commands executed by a low-level controller, thereby eliminating conventional contact estimation and control law design. Real-world experiments on cylindrical peg insertion demonstrate that incorporating tactile feedback increases the success rate from 50% to 75%, effectively validating the feasibility and superiority of the proposed framework.
📝 Abstract
Robots that autonomously determine their actions from language instructions and sensory observations could perform new contact-rich manipulation tasks without task-specific training or hand-designed rules. To perform these tasks, robots must infer how objects contact one another and move as a result, then select actions. For contact inference and action selection, prior approaches involve designing estimation models and tactile feedback control laws, or learning models for object-motion estimation, action-outcome prediction, and action selection from tactile data. Instead, we propose TacZero, which uses a pretrained general-purpose vision-language model (VLM) to interpret visual and tactile observations and select robot actions without additional tactile or manipulation training or task-specific rules for contact interpretation or action selection. TacZero provides the VLM with camera images, robot state, and three-axis tactile responses represented as numerical values or vectors overlaid on the images. From these observations and interaction history, the VLM generates commands specifying target end-effector positions and gripper opening or closing, which a low-level controller executes. In real-world cylindrical-peg insertion experiments, TacZero succeeded in 15 of 20 trials with numerical tactile input, compared with 10 of 20 without tactile input. This study provides a concrete starting point for further research on contact-rich manipulation using general-purpose VLMs and highlights challenges in pursuing this direction.