🤖 AI Summary
This study addresses the limitations of existing visuomotor policies that rely on global feature regression, which suffer from low data efficiency and poor viewpoint robustness. To overcome the constraints of conventional MLP-based regression, this work proposes substituting implicit action-feature learning with camera geometric priors. Specifically, the method leverages a pretrained vision encoder to extract 2D image features and binds them to the 3D action space via camera projection geometry, incorporating a discretized candidate scoring mechanism for precise action selection. The resulting architecture exhibits strong spatial consistency, achieving near-perfect task completion with only five demonstrations. Furthermore, it demonstrates exceptional generalization capabilities under extreme viewpoint shifts and previously unseen object configurations.
📝 Abstract
We introduce BIND, a new action representation for visuomotor robot policies that binds 3D robot actions to their corresponding 2D image features, yielding strong data efficiency gains and robustness to out-of-distribution object positions and camera viewpoints. The action heads of current robot policies are typically formulated as an MLP regression from a single global feature vector produced by a pre-trained vision encoder. This global formulation requires the policy network to discover, from demonstrations alone, the relationship between target robot actions and the image features they project onto. The consequence is that although modern image features are semantically descriptive, spatially robust, and even multiview-consistent, the policies built on them are brittle to subtle changes in camera viewpoint and object placement--and surprisingly data-inefficient. BIND closes this gap by supplying the action-feature relationship through camera geometry rather than learning: it discretizes a volume of candidate end effector positions, attaches each candidate to the pre-trained features at its projection in each camera view, and selects actions by scoring each candidate's position and image-bound feature combination. On a real robot, we study data efficiency and out-of-distribution robustness to unseen object positions and camera viewpoints, as well as general long-horizon task execution and dexterity. We find BIND to be highly data-efficient and robust: it achieves near-perfect success on tasks with as few as 5 demonstrations, and degrades gracefully under steep camera-viewpoint shifts and held-out object positions where coordinate-regression baselines completely fail.