Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing vision-language-action (VLA) models, which rely on limited fields of view and thus struggle to support mobile manipulation tasks requiring global spatial awareness, compounded by the scarcity of high-quality full-body demonstration data. To overcome these challenges, the authors develop a VR-based full-body teleoperation system to coordinate a wheeled dual-arm robot and collect 5.5 hours of real-world multimodal demonstration data. They introduce PanoVLA, the first VLA framework incorporating panoramic vision, featuring a dedicated panoramic encoder and a multimodal fusion module within a hybrid Transformer architecture to enable end-to-end, language-conditioned action generation. Evaluated on four real-world mobile manipulation tasks, PanoVLA achieves an average stage completion rate of 91.3% and an end-to-end success rate of 73.4%, substantially outperforming baseline methods using only local visual observations.
📝 Abstract
Mobile manipulation is a key capability for embodied intelligence, enabling robots to accomplish complex multi-stage tasks in open-world environments. However, mobile manipulation poses two key challenges for vision-language-action (VLA) policies: At the data level, the efficient collection of high-quality whole-body demonstrations demands the coordinated control of both the mobile base and the robotic arms; at the model level, existing VLA models predominantly rely on local camera observations, whose limited field of view hinders global spatial understanding. To address these challenges, we develop a whole-body teleoperation system and a panoramic-aware VLA policy. The system enables coordinated control of a wheeled bimanual robot through a single VR interface and supports the acquisition of a real-world mobile manipulation dataset comprising 5.5 hours of multimodal demonstrations. Building upon this dataset, we propose PanoVLA, a panorama-aware vision-language-action policy for mobile bimanual manipulation. Built upon a Mixture-of-Transformers architecture, PanoVLA introduces global spatial context through dedicated panorama encoding and fusion modules, enabling effective integration of panoramic observations with language instructions and robot states for action generation. Evaluation on four real-world mobile manipulation tasks demonstrates that PanoVLA achieves an average stage completion rate of 91.3\% and an end-to-end success rate of 73.4\%, substantially outperforming local-view baselines. These results demonstrate that incorporating panoramic spatial context improves spatial understanding and closed-loop manipulation performance in mobile robots.
Problem

Research questions and friction points this paper is trying to address.

mobile manipulation
vision-language-action
panoramic perception
whole-body teleoperation
spatial understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

panorama-aware
whole-body teleoperation
vision-language-action
mobile manipulation
Mixture-of-Transformers