🤖 AI Summary
This work proposes the Vision-Click-Action (VCA) framework to address the limitations of existing vision-language-action (VLA) models, which rely on linguistic instructions and often suffer from ambiguity and high cognitive load in scenarios with vague targets, hindering precise robotic manipulation. VCA replaces textual commands with click-based visual interaction, leveraging a pre-trained image segmentation model to establish an end-to-end mapping from 2D click coordinates to robot action control. This approach substantially reduces interpretation errors and operator cognitive burden, achieving higher accuracy in instance-level object selection and improved efficiency in sequential task execution. By enabling intuitive, non-linguistic human-robot interaction, VCA offers a scalable new paradigm for real-world applications.
📝 Abstract
The reliance on language in Vision-Language-Action (VLA) models introduces ambiguity, cognitive overhead, and difficulties in precise object identification and sequential task execution, particularly in environments with multiple visually similar objects. To address these limitations, we propose Vision-Click-Action (VCA), a framework that replaces verbose textual commands with direct, click-based visual interaction using pretrained segmentation models. By allowing operators to specify target objects clearly through visual selection in the robot's 2D camera view, VCA reduces interpretation errors, lowers cognitive load, and provides a practical and scalable alternative to language-driven interfaces for real-world robotic manipulation. Experimental results validate that the proposed VCA framework achieves effective instance-level manipulation of specified target objects. Experiment videos are available at https://robrosinc.github.io/vca/.