ExStereo: Lifting 2D Vision-Language-Action Models to 3D with Explicit Stereo Representations

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of Vision-Language-Action (VLA) models in high-precision robotic manipulation caused by their reliance on monocular RGB inputs, which lack depth perception. To overcome this, we propose a stereo vision enhancement framework that leverages a stereo matching foundation model to reconstruct 3D geometry. We design an action-stereo cross-attention mechanism that, combined with multi-view rendering and a self-supervised intermediate training stage, efficiently integrates explicit stereo features into the action generation process of 2D VLA models. Experiments conducted in both simulation and on a real-world bimanual platform demonstrate that fine-tuned π0.5 and SmolVLA significantly outperform baseline methods. These results validate the effectiveness of incorporating stereo perception for enhancing manipulation precision.
📝 Abstract
Three-dimensional perception is critical for robotic manipulation, particularly for high-precision tasks, as recovering metric depth and precise 3D object positions from monocular RGB observations is inherently ill-posed. However, many Vision-Language-Action (VLA) models rely solely on RGB observations for perception. Leveraging recent advances in foundation models for stereo matching, we introduce ExStereo, a stereo module that augments pre-trained 2D VLAs with 3D perception. ExStereo reconstructs scene geometry from stereo image pairs and renders multi-view observations as an explicit stereo representation for stereo feature extraction. The action tokens from the action expert selectively attend to the resulting stereo tokens through our proposed action-stereo cross-attention mechanism, enabling the policy to generate robot actions conditioned on 3D scene information. To learn robust 3D representations, we introduce a mid-training stage before task-specific post-training, using a self-supervised learning objective on large-scale stereo data. We validate our approach by fine-tuning two publicly available VLAs, $\pi_{0.5}$ and SmolVLA, and evaluate them in simulation and on a real-world bimanual PiPER platform. Across both settings, VLAs fine-tuned with ExStereo consistently outperform baselines, demonstrating the effectiveness of stereo perception for robotic manipulation. Our project website is at: https://exstereo-vla.github.io/ExStereo/.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
3D perception
robotic manipulation
stereo representation
monocular RGB
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Explicit Stereo Representations
Action-Stereo Cross-Attention
Self-Supervised Learning
Robotic Manipulation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
I
I-Chun Arthur Liu
Viterbi School of Engineering, University of Southern California
J
Jason Chen
Viterbi School of Engineering, University of Southern California
G
Gaurav S. Sukhatme
Viterbi School of Engineering, University of Southern California
Daniel Seita
Daniel Seita
University of Southern California
RoboticsMachine Learning