Score
Designs and implements map representations and algorithms that fuse spatial geometry with semantic labels and per-instance properties, producing context-aware and panoptic semantic maps. Builds systems to register and update semantics on SLAM/point-cloud/occupancy geometry, store instance classes and attributes, and support context-aware querying and map filtering.
To address the high memory overhead and computational inefficiency in semantic mapping for embodied agents operating long-term in unknown indoor environments, this paper presents the first comprehensive, multi-dimensional survey tailored to indoor scenarios. We systematically classify existing approaches along two orthogonal dimensions—structural representation (grid-based, topological, point-cloud, or hybrid) and information type (implicit features vs. explicit data)—and critically analyze state-of-the-art methods, including deep semantic segmentation, SLAM-semantic integration, cross-modal alignment, graph neural network (GNN)-based modeling, and 3D reconstruction, identifying their technical limits and bottlenecks. We propose an evolutionary paradigm for semantic maps characterized by open-vocabulary support, queryability, and task-agnosticism. Finally, we distill four key future research directions. This work establishes a theoretical framework and technical roadmap for advancing semantic understanding, long-horizon navigation, and task planning in embodied intelligence.
The semantic visual SLAM field lacks a systematic survey, particularly regarding the integration of deep learning and large language models (LLMs). To address this gap, we propose a unified problem formulation that decomposes semantic SLAM into five core modules: visual localization, semantic feature extraction, map construction, data association, and loop closure optimization. We introduce a modular analytical framework that unifies classical geometric approaches with modern semantic understanding techniques—including semantic segmentation, object detection, scene understanding, and LLM-based reasoning—and conduct empirical evaluations on benchmark datasets. Our work provides the first comprehensive taxonomy of technical evolution, critically analyzes limitations of existing methods, and identifies key bottlenecks: semantic consistency, cross-modal alignment, and real-time performance. The study establishes an authoritative knowledge base and a scalable technical roadmap for future research in semantic SLAM.
This work addresses the limitation of existing autonomous mobile robots in logistics scenarios, which rely solely on geometric maps and lack understanding of object-level semantics and contextual properties such as movability. To overcome this, the authors propose a context-aware semantic mapping approach that integrates SLAM, SAM-based instance segmentation, instance clustering, and multi-view visual language model (VLM) reasoning. Operating in a zero-shot, open-vocabulary setting without task-specific training, the method leverages two prompting strategies to coordinate three VLMs for multi-view semantic fusion. The system achieves a semantic classification mIoU of 98.93% and a movability estimation mean accuracy of 89.17%, substantially enhancing contextual awareness and navigation robustness in dynamic logistics environments.
To address the challenge of balancing multi-view observation quality and traversal redundancy in autonomous semantic exploration and dense semantic mapping for ground robots operating in complex, unknown environments, this paper proposes a decoupled hierarchical planning framework. Our method introduces three key contributions: (1) a novel decoupled priority-based local sampler that explicitly models multi-view semantic observation requirements; (2) a safety-aggressive dual-mode exploration state machine coupled with a voxel-level complete coverage strategy; and (3) a plug-and-play semantic mapping module enabling LiDAR-panoramic camera fusion perception and SLAM-integrated point-cloud-level semantic mapping. Evaluated in both simulation and real-world unstructured environments, the approach significantly improves exploration efficiency, reduces path length, and ensures high-accuracy dense semantic object reconstruction with comprehensive multi-view coverage.
Current semantic SLAM systems face limitations in semantic understanding depth, robustness of loop closure detection, and stability of data association. This work proposes an RGB-based semantic SLAM framework grounded in a unified geometry-instance representation, which for the first time extends geometric foundation models to jointly predict dense geometric structures and view-consistent instance embeddings. By introducing viewpoint-invariant semantic anchors, the method bridges the gap between geometric reconstruction and open-vocabulary semantic mapping, enabling semantic-coherent data association and instance-guided loop closure detection. Experimental results demonstrate that the proposed system significantly outperforms state-of-the-art approaches in terms of map consistency and reliability under large-baseline loop closure scenarios.
This work addresses language-query-based embodied vision-language navigation under open-set conditions. Methodologically, it introduces a 3D Open-Set Semantic Instance Mapping (O3D-SIM) framework that integrates multimodal foundation models—specifically CLIP for cross-modal alignment and zero-shot object recognition, and SAM for image-level segmentation—combined with SLAM-based pose estimation and point-cloud instance clustering to construct an open-set, instance-level, semantically enriched 3D map. The core contribution is the first realization of open-set 3D instance semantic mapping capable of supporting language queries involving previously unseen object categories, thereby overcoming the limitations of conventional closed-set semantic mapping. Experimental results demonstrate substantial improvements in language-guided navigation success rates; qualitative analysis further confirms strong generalization capability and interpretability of the generated maps.
Traditional panoramic semantic mapping is constrained by predefined categories and struggles to recognize unknown objects. To address this, we propose the first promptable unified panoramic mapping framework. Our method integrates natural language prompts with multimodal foundation models (CLIP and SAM) to construct a prompt-driven dynamic labeling module, enabling real-time open-vocabulary semantic parsing. By jointly leveraging 3D reconstruction and instance segmentation, the framework achieves end-to-end promptable panoramic mapping. Extensive evaluation on both real-world and synthetic datasets demonstrates significant improvements in unknown-object segmentation accuracy and semantic labeling fidelity, while enabling natural-language-guided interactive map construction. Ablation studies further validate that foundation-model-based dynamic labeling substantially outperforms conventional fixed-label paradigms. This work establishes a new paradigm for flexible, scalable, and user-controllable semantic mapping beyond closed-set assumptions.
该研究提出一种结合外部校准相机、对象检测、持续跟踪及本体驱动语义更新的混合管道,以构建动态语义世界模型,解决机器人在复杂环境中的交互问题。
Existing semantic mapping approaches struggle to maintain semantic consistency and robustness in dynamic environments due to high computational costs and the neglect of spatiotemporal relationships among voxels. This work proposes a spatiotemporally aware continuous semantic mapping framework that jointly models spatial structure and temporal consistency for the first time. By integrating voxel-level semantic reasoning, local uncertainty–driven adaptive inference ranges, and cross-frame semantic label fusion, the method significantly enhances both mapping efficiency and stability. Evaluated on SemanticKITTI, the approach achieves a mean Intersection over Union (mIoU) of 54.92%, representing a 13.18 percentage point improvement over purely spatial methods, along with approximately a 12% increase in mapping accuracy.
为解决机器人导航中开放词汇感知与环境长期变化的融合问题,SuperMap提出了一种结合高频几何SLAM和异步开放词汇感知的4D时空映射框架。
This work addresses the challenge of cross-view localization for robots operating in unseen urban environments using publicly available semantic maps such as OpenStreetMap. The proposed method leverages a vision-language model (VLM) to extract semantic landmarks from panoramic images and employs a lightweight matcher—distilled from the VLM via knowledge distillation—to efficiently align these observations with large-scale semantic maps. Temporal pose estimates are refined through Bayesian filtering. Trained solely on daytime data from a single city, the framework demonstrates robust generalization across eleven diverse environmental conditions, including nighttime and blizzard scenarios, and scales to map areas spanning hundreds of square kilometers. The authors also release a semantic dataset and source code to support reproducibility and further research.
This work addresses the challenge of robust drone localization in GNSS-denied environments, where visual-inertial odometry suffers from cumulative drift and appearance-based methods are sensitive to seasonal changes, lighting variations, viewpoint shifts, and map obsolescence. To overcome these limitations, the authors propose a semantic map localization framework that integrates structured semantic geometry with a stability-aware mechanism, leveraging persistent environmental structures such as roads and buildings. The approach employs semantic grid alignment, relational graph matching, geospatial saliency evaluation, and a ternary observation model (positive, contradictory, or unknown), complemented by an integrity-aware fuzzy rejection strategy. Experimental results demonstrate a Recall@1 of 94.5–95.5% across 220 cross-view queries, substantially outperforming global semantic descriptors (58.6%) and confirming the method’s robustness to perturbations including rotation, scaling, and occlusion.