PanoFuse: Panorama-Enhanced Vision-Language-Action Learning with Decoupled Semantic-Geometric Routing

๐Ÿ“… 2026-09-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the limited operational robustness of existing Vision-Language-Action (VLA) systems in occluded or complex environments, which stems from global context deficiency due to restricted fields of view. To overcome this, we propose a panoramic-enhanced VLA framework that leverages a pretrained panoramic foundation model to extract complementary semantic and geometric representations. Furthermore, we introduce a Decoupled Semantic-Geometric Routing (DSGR) mechanism that selectively fuses contextual streams via structured block attention, injecting global spatial information while preserving task-relevant semantics. A synchronized data collection pipeline and a corresponding real-world dataset are also constructed. Experimental results demonstrate that the proposed method achieves an average success rate of 52.9% across seven scenarios, significantly outperforming baselines under challenging conditions involving novel objects, unseen backgrounds, and strong distractors.
๐Ÿ“ Abstract
Vision-Language-Action (VLA) policies have shown promising performance in language-conditioned robotic manipulation. However, most existing VLA systems rely on conventional perspective cameras with limited fields of view, often missing global scene context and leading to unreliable manipulation under visual occlusions, distractors, and unseen environments. In this work, we propose PanoFuse, a panorama-enhanced VLA framework that complements local manipulation observations with global panoramic perception. PanoFuse introduces a dedicated panoramic branch that leverages a pretrained panoramic foundation model to extract complementary semantic and geometric representations from omnidirectional observations. Rather than directly mixing these heterogeneous features, we introduce Decoupled Semantic-Geometric Routing (DSGR), which maintains semantic and geometric representations as separate context streams and selectively routes both to downstream state and action representations through structured block-wise attention. This design provides the action expert with global spatial context while preserving task-relevant semantic information from the pretrained VLA backbone. We further develop a synchronized data collection pipeline and construct a new real-world manipulation dataset containing panoramic RGB observations, wrist-view images, language instructions, robot states, and actions. Across seven evaluation settings, PanoFuse achieves an average success rate of 52.9%, outperforming the evaluated baselines and achieving consistent gains under novel-object, unseen-background, and distractor-rich settings. Code and data will be released publicly at https://xux-hnu.github.io/PanoFuse.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
panoramic perception
limited field of view
visual occlusion
robotic manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
Panoramic Perception
Decoupled Semantic-Geometric Routing
Robotic Manipulation
Block-wise Attention
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
P
Peng Xu
School of Artificial Intelligence and Robotics, Hunan University, China
H
Haoran Lin
School of Artificial Intelligence and Robotics, Hunan University, China
W
Wanjun Jia
School of Artificial Intelligence and Robotics, Hunan University, China
K
Kai Luo
School of Artificial Intelligence and Robotics, Hunan University, China
Wenrui Chen
Wenrui Chen
Hunan University
RoboticsHandsGraspingDexterous ManipulationHuman-Robot Collaboration
Zhiyong Li
Zhiyong Li
Professor of Computer Science, Hunan University
computer vision๏ผŒobject detection
Kailun Yang
Kailun Yang
Professor. School of Artificial Intelligence and Robotics, Hunan University (HNU); KIT; UAH; ZJU
Computer VisionComputational OpticsIntelligent VehiclesAutonomous DrivingRobotics