🤖 AI Summary
This study addresses the sensitivity of visual imitation learning policies to viewpoint perturbations, which hinders their deployment across diverse environments. Through controlled empirical investigations, we systematically identify key design factors that enhance the cross-viewpoint generalization of visuomotor policies. Our findings reveal that preserving dense visual tokens and guiding the action head to engage in geometric reasoning significantly improve viewpoint robustness. Evaluated via multi-camera pose simulation, the proposed approach achieves zero-shot sim-to-real transfer while maintaining high performance under randomized camera configurations. These results offer critical insights for the reliable cross-viewpoint deployment of visuomotor policies in real-world settings.
📝 Abstract
Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations, making deployment in diverse environments a challenge. We present a controlled empirical study of which design choices allow visuomotor policies to generalize across viewpoints. We find that viewpoint generalization improves when dense visual tokens are retained and the action head participates in geometric reasoning. On a suite of simulated tasks that span a wide range of camera poses, we show that these design choices yield a policy that remains performant across viewpoints. As a practical consequence, a policy trained with these design choices also transfers zero-shot from simulation to the real world under random camera configurations.