🤖 AI Summary
This study addresses the limitations of restricted local coverage and the absence of persistent open-vocabulary semantic occupancy mapping in panoramic perception. To overcome these challenges, this work proposes a training-free framework that unifies geometric and semantic evidence through a novel online architecture. By integrating panoramic SLAM, open-vocabulary perception, and a long-term voxel memory mechanism, the proposed method constructs globally consistent, language-queryable semantic maps. Furthermore, two new benchmarks, Pan-Replica and Pan-Holo360D, are established to facilitate evaluation. Experimental results demonstrate that the proposed approach yields substantial improvements in Intersection over Union (IoU) metrics across both benchmarks, achieving absolute gains of 20.03/7.06 and 43.26/20.16, respectively.
📝 Abstract
Persistent semantic occupancy mapping is essential for embodied scene understanding. However, perspective-based systems provide limited spatial coverage, while existing panoramic methods primarily predict local volumes from single observations. We introduce PanOVOcc, a training-free framework for persistent open-vocabulary semantic occupancy mapping from panoramic sequences. PanOVOcc unifies panoramic SLAM, open-vocabulary perception, and long-term spatial voxel memory within an online architecture, continuously integrating geometric and semantic evidence into a global, language-queryable map. To facilitate systematic evaluation of this setting, we establish Pan-Replica and Pan-Holo360D, two benchmarks pairing continuous panoramic RGB-D sequences with scene-level semantic occupancy ground truth across synthetic and real-world scenes. Compared with the strongest evaluated baseline for each metric, PanOVOcc improves occupancy IoU and semantic mIoU by absolute +20.03 and +7.06 on Pan-Replica, and by +43.26 and +20.16 on Pan-Holo360D, respectively. The source code and the established benchmarks will be available at https://github.com/bakereet/PanOVOcc.