Score
Designs, builds, and evaluates systems that enable an embodied agent to simultaneously localize itself and build a spatial map (SLAM) while actively selecting actions to explore, reduce uncertainty, and fuse sensor data into robust geometric maps. Extends that work by incorporating contextual and semantic reasoning—adding semantic labeling, egocentric or high‑level model outputs, and policies that balance frontier exploration with goal‑directed semantic navigation to ground abstract information in a 3D map.
The semantic visual SLAM field lacks a systematic survey, particularly regarding the integration of deep learning and large language models (LLMs). To address this gap, we propose a unified problem formulation that decomposes semantic SLAM into five core modules: visual localization, semantic feature extraction, map construction, data association, and loop closure optimization. We introduce a modular analytical framework that unifies classical geometric approaches with modern semantic understanding techniques—including semantic segmentation, object detection, scene understanding, and LLM-based reasoning—and conduct empirical evaluations on benchmark datasets. Our work provides the first comprehensive taxonomy of technical evolution, critically analyzes limitations of existing methods, and identifies key bottlenecks: semantic consistency, cross-modal alignment, and real-time performance. The study establishes an authoritative knowledge base and a scalable technical roadmap for future research in semantic SLAM.
To address the high memory overhead and computational inefficiency in semantic mapping for embodied agents operating long-term in unknown indoor environments, this paper presents the first comprehensive, multi-dimensional survey tailored to indoor scenarios. We systematically classify existing approaches along two orthogonal dimensions—structural representation (grid-based, topological, point-cloud, or hybrid) and information type (implicit features vs. explicit data)—and critically analyze state-of-the-art methods, including deep semantic segmentation, SLAM-semantic integration, cross-modal alignment, graph neural network (GNN)-based modeling, and 3D reconstruction, identifying their technical limits and bottlenecks. We propose an evolutionary paradigm for semantic maps characterized by open-vocabulary support, queryability, and task-agnosticism. Finally, we distill four key future research directions. This work establishes a theoretical framework and technical roadmap for advancing semantic understanding, long-horizon navigation, and task planning in embodied intelligence.
To address low exploration efficiency and inaccurate global semantic map construction in unknown environments, this paper proposes a hierarchical exploration framework leveraging semantic map prediction. The method integrates a long-term environmental understanding mechanism with a reinforcement learning–driven reward function, iteratively predicting semantic distributions in unobserved regions and guiding exploration path planning via prediction–ground-truth discrepancy. A hierarchical decision-making architecture is further introduced to optimize long-horizon exploration policies. Experiments on standard benchmarks demonstrate that, under identical time budgets, the proposed approach significantly improves map coverage (+12.7%) and semantic mapping accuracy (mIoU +8.3%) over current state-of-the-art methods.
This work addresses the autonomous exploration problem for mobile robots operating in unknown environments, where geometric and semantic mapping must be performed simultaneously. We propose the first semantic-guided next-best-view (NBV) selection framework. Our method formalizes “semantic exploration” as a novel task, introduces a semantic visibility scoring mechanism to enable active perception jointly optimized for structural and semantic map construction, and integrates semantic segmentation networks with 3D reconstruction and multi-view sampling optimization. Evaluated in both simulation and on real robotic platforms, the approach significantly improves semantic map accuracy and environmental understanding. Moreover, it enhances generalization performance on downstream tasks—including object localization and scene-based question answering—by leveraging semantically informed viewpoint selection. The framework bridges the gap between traditional geometry-driven NBV strategies and high-level semantic reasoning, enabling more intelligent and task-aware robotic exploration.
To address the challenge of autonomous robotic exploration in unknown environments—requiring simultaneous high-fidelity geometric reconstruction and robust semantic understanding—this paper proposes ActiveSGM, a novel active mapping framework. ActiveSGM introduces the first semantic uncertainty quantification mechanism built upon 3D Gaussian Splatting, tightly coupling sparse semantic encoding with geometric uncertainty modeling to enable joint semantic-geometric optimization. By predicting information gain from candidate viewpoints, it dynamically selects optimal observations, closing the “explore–perceive–map” loop. Experiments on Replica and Matterport3D demonstrate that ActiveSGM significantly improves mapping completeness (+18.7%), semantic segmentation accuracy (mIoU +12.3%), and robustness to sensor noise. The framework establishes a new paradigm for adaptive, open-world autonomous exploration.
This work addresses the challenge of autonomous exploration and mapping for size-, weight-, and power-constrained (SWaP-limited) UAVs in multi-floor, GPS-denied indoor environments. We propose a metric-semantic joint active SLAM framework. Our key contributions are: (1) the first integration of semantic loop closure (SLC) into an active SLAM policy, enabling synergistic optimization between exploration behavior and pose uncertainty reduction; and (2) a lightweight algorithm based on sparse information abstraction, specifically designed to comply with onboard computational constraints. The system jointly performs metric mapping, semantic recognition, loop closure correction, and online exploration decision-making. Experimental results demonstrate median translational and yaw errors reduced by 90% and 75%, respectively, while pose uncertainty and semantic map uncertainty decrease by 70% and 65%. These improvements significantly enhance both mapping accuracy and exploration efficiency in complex indoor settings.
This study addresses the fragmented treatment of agent architectures and spatial intelligence in existing research, which lacks a unified framework integrating perception, reasoning, and physical action—thereby limiting the effectiveness of embodied agents in real-world 3D environments. Through a systematic review of over 2,000 papers, this work proposes the first triaxial taxonomy that explicitly distinguishes spatial embodiment (geometric and physical) from symbolic embodiment, and constructs an analytical framework combining graph neural networks (GNNs), large language models (LLMs), and world models. The research highlights the critical roles of hierarchical memory, GNN–LLM synergy, and world models in cross-scale spatial tasks, yielding three core insights and identifying six key challenges. These contributions establish a standardized evaluation benchmark and chart a roadmap for future advancements in robotics, autonomous driving, and geospatial intelligence.
为解决LLM与ROS导航集成问题,提出一种通过MCP协议连接LLM推理与ROS导航的框架,实现自然语言指令下的高效自主导航。
研究提出OptiSight框架,结合语义理解和几何控制解决室内自主导航问题,通过视觉-语言模型和几何伺服实现高效导航。
This work addresses the challenge of balancing semantic understanding and efficient exploration for micro aerial vehicles (MAVs) in complex, unstructured 3D environments, which critically impacts search-and-rescue performance. The authors propose a semantic-guided viewpoint planning framework that tightly integrates semantic reasoning with 3D exploration. Leveraging a large language model (LLM), the method generates semantic priors to assess target similarity and propagates semantic priorities to frontier voxels through active perception, computing semantic information gain to guide viewpoint selection. A compositional planner then produces efficient exploration trajectories. Experimental results in simulation demonstrate significant improvements over baseline approaches in rapidly locating targets while controlling exploration time. Real-world MAV trials further validate the framework’s practicality under constraints of limited battery life, narrow perceptual range, and semantic uncertainty.
This work addresses the limitations of semantic navigation—namely, its reliance on partial observations, greedy decision-making, and poor efficiency over long distances—stemming from the absence of a structured global representation. To overcome these challenges, the authors propose a zero-shot navigation method based on a hierarchical 3D scene graph (HSG). The approach constructs a multi-granularity semantic topology online and, for the first time, employs HSG as an abstract global state representation. Integrated with a belief-driven hierarchical planning mechanism, it combines semantic priors with exploration evidence to simulate macro-actions and evaluate their long-term returns under limited visibility. Experiments in multiple high-fidelity simulation environments demonstrate that the method significantly outperforms current state-of-the-art approaches, achieving average improvements of 9.4% in success rate (SR) and 5.0% in success weighted by path length (SPL) on long-range tasks.