OC-VLA++: Monocular Geometry-Guided Cross-View Consistency for Viewpoint-Robust Robotic Manipulation

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the significant performance degradation of existing robotic manipulation policies under limited camera viewpoint coverage, where generalization to unseen viewpoints remains challenging. To overcome this limitation, the authors propose a geometry-guided, cross-view action-equivariant learning framework that explicitly models the geometric transformation relationships between action predictions across different viewpoints. By integrating monocular scene geometry, camera-frame action grounding, and pairwise viewpoint supervision, the approach transcends conventional reliance on image augmentations or single-view supervision. Experimental results demonstrate that the proposed method substantially improves task success rates under novel viewpoints and exhibits more graceful performance degradation as camera displacement increases, thereby achieving enhanced viewpoint robustness in robotic manipulation.
📝 Abstract
We propose OC-VLA++, an extension of OC-VLA for viewpoint generalization under limited camera coverage. While OC-VLA grounds robot actions in the camera coordinate system to align action supervision with visual observations, camera-space grounding alone can still overfit to the few viewpoints observed during training. OC-VLA++ addresses this limitation by introducing geometry-guided paired-view supervision and an explicit cross-view action-equivariance objective. Given paired observations of the same manipulation scene from geometrically related viewpoints, the model is trained such that their camera-space predictions correspond to the same robot-frame action. This objective explicitly supervises how action predictions should transform across viewpoints, rather than relying solely on image-level augmentation. Experiments demonstrate substantial improvements in unseen-view generalization under limited camera coverage, with performance degrading more gracefully under increasing camera displacement. These results establish cross-view action equivariance as an effective complement to observation-centric action grounding for robust real-world deployment.
Problem

Research questions and friction points this paper is trying to address.

viewpoint generalization
robotic manipulation
camera coverage
cross-view consistency
action grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-view consistency
action equivariance
geometry-guided supervision
viewpoint generalization
monocular manipulation
🔎 Similar Papers
No similar papers found.