contextual active slam

Designs, builds, and evaluates systems that enable an embodied agent to simultaneously localize itself and build a spatial map (SLAM) while actively selecting actions to explore, reduce uncertainty, and fuse sensor data into robust geometric maps. Extends that work by incorporating contextual and semantic reasoning—adding semantic labeling, egocentric or high‑level model outputs, and policies that balance frontier exploration with goal‑directed semantic navigation to ground abstract information in a 3D map.

contextualactiveslam

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Semantic Mapping in Indoor Embodied AI -- A Comprehensive Survey and Future Directions

Jan 10, 2025
SR
Sonia Raychaudhuri
🏛️ Simon Fraser University

To address the high memory overhead and computational inefficiency in semantic mapping for embodied agents operating long-term in unknown indoor environments, this paper presents the first comprehensive, multi-dimensional survey tailored to indoor scenarios. We systematically classify existing approaches along two orthogonal dimensions—structural representation (grid-based, topological, point-cloud, or hybrid) and information type (implicit features vs. explicit data)—and critically analyze state-of-the-art methods, including deep semantic segmentation, SLAM-semantic integration, cross-modal alignment, graph neural network (GNN)-based modeling, and 3D reconstruction, identifying their technical limits and bottlenecks. We propose an evolutionary paradigm for semantic maps characterized by open-vocabulary support, queryability, and task-agnosticism. Finally, we distill four key future research directions. This work establishes a theoretical framework and technical roadmap for advancing semantic understanding, long-horizon navigation, and task planning in embodied intelligence.

Computational EfficiencyMemory EfficiencySemantic Mapping

SEA: Semantic Map Prediction for Active Exploration of Uncertain Areas

Oct 22, 2025
HD
Hongyu Ding
🏛️ Nanjing University | Cardiff University

To address low exploration efficiency and inaccurate global semantic map construction in unknown environments, this paper proposes a hierarchical exploration framework leveraging semantic map prediction. The method integrates a long-term environmental understanding mechanism with a reinforcement learning–driven reward function, iteratively predicting semantic distributions in unobserved regions and guiding exploration path planning via prediction–ground-truth discrepancy. A hierarchical decision-making architecture is further introduced to optimize long-horizon exploration policies. Experiments on standard benchmarks demonstrate that, under identical time budgets, the proposed approach significantly improves map coverage (+12.7%) and semantic mapping accuracy (mIoU +8.3%) over current state-of-the-art methods.

Improves global map coverage through semantic prediction and active explorationPredicts missing map areas to guide robot exploration efficientlyUses reinforcement learning for long-term exploration strategy optimization

SeGuE: Semantic Guided Exploration for Mobile Robots

Apr 04, 2025
CS
Cody Simons
🏛️ University of California, Riverside

This work addresses the autonomous exploration problem for mobile robots operating in unknown environments, where geometric and semantic mapping must be performed simultaneously. We propose the first semantic-guided next-best-view (NBV) selection framework. Our method formalizes “semantic exploration” as a novel task, introduces a semantic visibility scoring mechanism to enable active perception jointly optimized for structural and semantic map construction, and integrates semantic segmentation networks with 3D reconstruction and multi-view sampling optimization. Evaluated in both simulation and on real robotic platforms, the approach significantly improves semantic map accuracy and environmental understanding. Moreover, it enhances generalization performance on downstream tasks—including object localization and scene-based question answering—by leveraging semantically informed viewpoint selection. The framework bridges the gap between traditional geometry-driven NBV strategies and high-level semantic reasoning, enabling more intelligent and task-aware robotic exploration.

Autonomous exploration for semantic and geometric mappingEnabling robots to better understand environmentsNext-best-view scoring based on semantic features

Understanding while Exploring: Semantics-driven Active Mapping

May 30, 2025
LC
Liyan Chen
🏛️ Stevens Institute of Technology | Goertek Alpha Labs | Purdue University

To address the challenge of autonomous robotic exploration in unknown environments—requiring simultaneous high-fidelity geometric reconstruction and robust semantic understanding—this paper proposes ActiveSGM, a novel active mapping framework. ActiveSGM introduces the first semantic uncertainty quantification mechanism built upon 3D Gaussian Splatting, tightly coupling sparse semantic encoding with geometric uncertainty modeling to enable joint semantic-geometric optimization. By predicting information gain from candidate viewpoints, it dynamically selects optimal observations, closing the “explore–perceive–map” loop. Experiments on Replica and Matterport3D demonstrate that ActiveSGM significantly improves mapping completeness (+18.7%), semantic segmentation accuracy (mIoU +12.3%), and robustness to sensor noise. The framework establishes a new paradigm for adaptive, open-world autonomous exploration.

Enhances robotic autonomy in unknown environments through proactive explorationImproves mapping completeness and accuracy with strategic viewpoint selectionPredicts informativeness of observations using semantic and geometric uncertainty

3D Active Metric-Semantic SLAM

Sep 13, 2023
YT
Yuezhan Tao
🏛️ University of Pennsylvania

This work addresses the challenge of autonomous exploration and mapping for size-, weight-, and power-constrained (SWaP-limited) UAVs in multi-floor, GPS-denied indoor environments. We propose a metric-semantic joint active SLAM framework. Our key contributions are: (1) the first integration of semantic loop closure (SLC) into an active SLAM policy, enabling synergistic optimization between exploration behavior and pose uncertainty reduction; and (2) a lightweight algorithm based on sparse information abstraction, specifically designed to comply with onboard computational constraints. The system jointly performs metric mapping, semantic recognition, loop closure correction, and online exploration decision-making. Experimental results demonstrate median translational and yaw errors reduced by 90% and 75%, respectively, while pose uncertainty and semantic map uncertainty decrease by 70% and 65%. These improvements significantly enhance both mapping accuracy and exploration efficiency in complex indoor settings.

Balances exploration efficiency and localization error reductionExplores GPS-denied indoor multi-floor environments with aerial robotsImproves robot pose and map accuracy via semantic loop closure

Latest Papers

What's happening recently
View more

This study addresses the fragmented treatment of agent architectures and spatial intelligence in existing research, which lacks a unified framework integrating perception, reasoning, and physical action—thereby limiting the effectiveness of embodied agents in real-world 3D environments. Through a systematic review of over 2,000 papers, this work proposes the first triaxial taxonomy that explicitly distinguishes spatial embodiment (geometric and physical) from symbolic embodiment, and constructs an analytical framework combining graph neural networks (GNNs), large language models (LLMs), and world models. The research highlights the critical roles of hierarchical memory, GNN–LLM synergy, and world models in cross-scale spatial tasks, yielding three core insights and identifying six key challenges. These contributions establish a standardized evaluation benchmark and chart a roadmap for future advancements in robotics, autonomous driving, and geospatial intelligence.

agentic reasoningembodied agentsspatial grounding

This work addresses the challenge of balancing semantic understanding and efficient exploration for micro aerial vehicles (MAVs) in complex, unstructured 3D environments, which critically impacts search-and-rescue performance. The authors propose a semantic-guided viewpoint planning framework that tightly integrates semantic reasoning with 3D exploration. Leveraging a large language model (LLM), the method generates semantic priors to assess target similarity and propagates semantic priorities to frontier voxels through active perception, computing semantic information gain to guide viewpoint selection. A compositional planner then produces efficient exploration trajectories. Experimental results in simulation demonstrate significant improvements over baseline approaches in rapidly locating targets while controlling exploration time. Real-world MAV trials further validate the framework’s practicality under constraints of limited battery life, narrow perceptual range, and semantic uncertainty.

3D explorationautonomous navigationcluttered environments

This work addresses the limitations of semantic navigation—namely, its reliance on partial observations, greedy decision-making, and poor efficiency over long distances—stemming from the absence of a structured global representation. To overcome these challenges, the authors propose a zero-shot navigation method based on a hierarchical 3D scene graph (HSG). The approach constructs a multi-granularity semantic topology online and, for the first time, employs HSG as an abstract global state representation. Integrated with a belief-driven hierarchical planning mechanism, it combines semantic priors with exploration evidence to simulate macro-actions and evaluate their long-term returns under limited visibility. Experiments in multiple high-fidelity simulation environments demonstrate that the method significantly outperforms current state-of-the-art approaches, achieving average improvements of 9.4% in success rate (SR) and 5.0% in success weighted by path length (SPL) on long-range tasks.

embodied agentsglobal representationhierarchical scene graph

Hot Scholars

PN

Peer Neubert

University of Koblenz
Autonomous SystemsArtificial IntelligenceComputer VisionMachine Learning
ER

Esa Rahtu

Professor, Tampere University, Finland
Computer VisionImage UnderstandingMachine Learning