๐ค AI Summary
This study addresses the limited operational robustness of existing Vision-Language-Action (VLA) systems in occluded or complex environments, which stems from global context deficiency due to restricted fields of view. To overcome this, we propose a panoramic-enhanced VLA framework that leverages a pretrained panoramic foundation model to extract complementary semantic and geometric representations. Furthermore, we introduce a Decoupled Semantic-Geometric Routing (DSGR) mechanism that selectively fuses contextual streams via structured block attention, injecting global spatial information while preserving task-relevant semantics. A synchronized data collection pipeline and a corresponding real-world dataset are also constructed. Experimental results demonstrate that the proposed method achieves an average success rate of 52.9% across seven scenarios, significantly outperforming baselines under challenging conditions involving novel objects, unseen backgrounds, and strong distractors.
๐ Abstract
Vision-Language-Action (VLA) policies have shown promising performance in language-conditioned robotic manipulation. However, most existing VLA systems rely on conventional perspective cameras with limited fields of view, often missing global scene context and leading to unreliable manipulation under visual occlusions, distractors, and unseen environments. In this work, we propose PanoFuse, a panorama-enhanced VLA framework that complements local manipulation observations with global panoramic perception. PanoFuse introduces a dedicated panoramic branch that leverages a pretrained panoramic foundation model to extract complementary semantic and geometric representations from omnidirectional observations. Rather than directly mixing these heterogeneous features, we introduce Decoupled Semantic-Geometric Routing (DSGR), which maintains semantic and geometric representations as separate context streams and selectively routes both to downstream state and action representations through structured block-wise attention. This design provides the action expert with global spatial context while preserving task-relevant semantic information from the pretrained VLA backbone. We further develop a synchronized data collection pipeline and construct a new real-world manipulation dataset containing panoramic RGB observations, wrist-view images, language instructions, robot states, and actions. Across seven evaluation settings, PanoFuse achieves an average success rate of 52.9%, outperforming the evaluated baselines and achieving consistent gains under novel-object, unseen-background, and distractor-rich settings. Code and data will be released publicly at https://xux-hnu.github.io/PanoFuse.