🤖 AI Summary
This study addresses the deployment vulnerability of Vision-Language-Action (VLA) models caused by fixed camera configurations by proposing a configurable camera framework. The method introduces a unified scene token interface to decouple scene representation from action learning, and achieves 3D-sensor-free viewpoint diversity through wrist-augmented pose sampling. By integrating multi-signal target view prediction, a lightweight spatial encoder, and a pretrained VLA backbone, the framework enables robust manipulation under arbitrary viewpoints. Evaluations on the RoboTwin benchmark and real-world robot experiments demonstrate that the proposed approach significantly outperforms existing methods, flexibly adapting to varying numbers of cameras and diverse pose configurations.
📝 Abstract
Vision-Language-Action (VLA) models have emerged as powerful foundations for robotic manipulation, but their reliance on fixed camera configurations during training makes them brittle to changes in camera count or pose during deployment. To overcome these limitations, we propose VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning. VersaCamVLA learns a unified scene-token interface that maps an arbitrary, variable set of posed RGB views into fixed-size latent scene tokens. This is achieved via multi-signal target-view prediction and Wrist-Augmented Pose Sampling (WAPS), which leverages natural wrist-camera motion for free pose diversity. At deployment, a lightweight spatial encoder injects these compact scene tokens into a pretrained base VLA as a supplementary visual condition, requiring no explicit 3D sensing or novel-view rendering. Experiments on RoboTwin, LIBERO, and a real-robot platform demonstrate that VersaCamVLA consistently outperforms prior VLA methods and direct multi-view baselines, maintaining robust performance across varying camera counts and unseen camera poses.