🤖 AI Summary
Vision-Language-Action (VLA) models are highly sensitive to wrist-camera pose variations, hindering cross-device deployment. This work proposes WARP-VLA, which employs a Mixture-of-Experts architecture with implicit viewpoint routing combined with expert networks to learn viewpoint-specific feature transformations. This approach effectively mitigates geometric cue distortions caused by minor mounting discrepancies, enabling robust policy execution without external calibration parameters. Evaluated on the LIBERO benchmark, WARP-VLA improves the success rate under wrist-view perturbations from 39.2% to 78.3%. Furthermore, the feature adaptation capability acquired through simulation training successfully transfers to diverse real-world robotic scenarios.
📝 Abstract
Despite recent advances in Vision-Language-Action models (VLAs) for robotic manipulation, their performance remains sensitive to changes in camera configuration. The problem becomes more evident in cross-setup deployment, as reproducing the exact camera pose used for training is nearly impossible. Unlike fixed external views, wrist views are more challenging because the camera moves with the robot, causing even small mounting variations to alter fine-grained geometric cues. To address this, we propose WARP-VLA, a camera-view robust VLA for diverse wrist camera configurations. WARP-VLA adopts a Mixture-of-Experts (MoE) architecture where individual experts learn view-specific feature transformations, and a router combines them based on implicit view information. This allows the policy to be deployed without requiring camera extrinsic parameters as additional input. Through experiments on the LIBERO benchmark, WARP-VLA improves the average success rate of pi-0.5 from 39.2% to 78.3% under wrist-view perturbations. The real-robot experiments further show that the feature-level adaptation learned in simulation successfully transfers to diverse deployment settings. To facilitate reproducibility and future research, we release our wrist viewpoint robustness benchmark and a plug-and-play implementation.