scene graph construction

Building structured spatial-semantic maps that encode objects, relationships, and geometry to support tasks like multi-session mapping, inferring support/order constraints, and compact storage for relocalization.

scenegraphconstruction

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing semantic mapping approaches struggle to maintain semantic consistency and robustness in dynamic environments due to high computational costs and the neglect of spatiotemporal relationships among voxels. This work proposes a spatiotemporally aware continuous semantic mapping framework that jointly models spatial structure and temporal consistency for the first time. By integrating voxel-level semantic reasoning, local uncertainty–driven adaptive inference ranges, and cross-frame semantic label fusion, the method significantly enhances both mapping efficiency and stability. Evaluated on SemanticKITTI, the approach achieves a mean Intersection over Union (mIoU) of 54.92%, representing a 13.18 percentage point improvement over purely spatial methods, along with approximately a 12% increase in mapping accuracy.

computational efficiencycontinuous semantic mappingdynamic scenes

Existing multimodal large language models lack explicit 3D geometric representations, limiting their ability to perform precise spatial understanding and reasoning. To address this, this work proposes an end-to-end 3D cognitive mapping framework that introduces, for the first time, an explicit 3D memory mechanism. By leveraging multi-view geometric modeling, the method anchors visual-semantic tokens into a unified 3D space, constructing a structured 3D cognitive map upon which the language model can directly conduct spatial reasoning. This approach effectively integrates semantic and geometric information, achieving state-of-the-art performance across multiple spatial reasoning benchmarks and significantly enhancing the model’s capacity for 3D scene understanding and question answering.

3D reasoninggeometric groundingmulti-view images

To bridge the gap between high-level language instructions and robotic execution, this work proposes a queryable 3D scene representation framework that unifies semantic interpretability with geometric fidelity. Methodologically, it introduces the first object-centric multimodal fusion architecture, integrating panoramic 3D reconstruction, high-precision point-cloud geometry modeling, structured 3D scene graphs, and semantic embeddings from large vision-language models (VLMs) to enable cross-modal object retrieval and higher-order spatial-semantic joint reasoning. Its core contribution is the first tight integration of queryable semantic representations with millimeter-accurate 3D geometry, supporting task planning from abstract natural-language directives. Evaluated in Unity simulation and a real-world wet-lab digital twin, the system accurately parses complex linguistic instructions and generates executable action sequences, significantly enhancing robotic high-level task performance in dynamic, cluttered environments.

Achieving comprehensive 3D scene understanding with semantic reasoning capabilitiesEnabling robots to comprehend high-level human instructions for complex tasksTranslating abstract language instructions into precise robotic task planning

Current vision-language models are constrained by limited context lengths, making it challenging to maintain temporally consistent spatial understanding without architectural modifications or fine-tuning. This work proposes a plug-and-play multi-agent framework that enables collaboration between local and global agents to construct a structured cognitive map as an external spatial memory. The approach requires no training or model architecture changes and introduces atomic map updates alongside cross-agent verification mechanisms, allowing seamless integration with any pretrained multimodal large language model. Experimental results demonstrate that the proposed framework significantly outperforms existing methods across multiple spatial reasoning benchmarks, effectively validating its efficacy and broad applicability.

cognitive mapcontext length limitationmodel-agnostic

This work addresses the limitation of traditional semantic mapping, which relies on predefined object categories and struggles to handle unknown objects, thereby hindering goal-directed exploration in partially unknown environments. To overcome this, the authors propose an open-vocabulary semantic mapping approach that integrates vision-language foundation models—such as CLIP—with 3D semantic scene graphs. This method introduces open-vocabulary semantic embeddings into 3D scene graph construction for the first time, effectively bypassing the constraints of fixed taxonomies. By enabling natural language–guided exploration strategies, the framework facilitates robust recognition and semantic reasoning about out-of-distribution target objects, significantly enhancing the robot’s semantic understanding and task generalization capabilities in dynamic and unfamiliar settings.

open-set recognitionout-of-distribution objectsrobotic exploration

Latest Papers

What's happening recently
View more

This work addresses the challenge of unified modeling for multi-class heterogeneous geographic entities in remote sensing vector mapping, where existing methods struggle to adequately represent topological relationships and instance boundaries. The authors propose reframing vector map construction as a structured text generation task by designing a GeoJSON-like hierarchical vector language that jointly encodes geometric, semantic, and topological information. A progressive vision-to-language mapping framework is introduced, optimized via reinforcement learning to ensure syntactic validity, content fidelity, and map executability of the generated output, thereby enabling cross-category unified modeling. Experiments on the newly curated VecMap-Bench dataset—comprising 54K images and 800K instances—demonstrate that the proposed approach significantly outperforms state-of-the-art methods in single- and multi-class mapping, cross-dataset transfer, and open-vocabulary generalization.

heterogeneous entity structuresinstance boundariesremote sensing vector mapping

This work addresses the limitation of existing autonomous mobile robots in logistics scenarios, which rely solely on geometric maps and lack understanding of object-level semantics and contextual properties such as movability. To overcome this, the authors propose a context-aware semantic mapping approach that integrates SLAM, SAM-based instance segmentation, instance clustering, and multi-view visual language model (VLM) reasoning. Operating in a zero-shot, open-vocabulary setting without task-specific training, the method leverages two prompting strategies to coordinate three VLMs for multi-view semantic fusion. The system achieves a semantic classification mIoU of 98.93% and a movability estimation mean accuracy of 89.17%, substantially enhancing contextual awareness and navigation robustness in dynamic logistics environments.

contextual semantic mappingintralogisticsobject movability

This work addresses the limitations of existing vision-language geolocation methods, which rely on point-to-point alignment and struggle to effectively model the semantic subspace jointly defined by images and text. To overcome this, the paper formulates the task as a multi-anchor geometric alignment problem and introduces a novel Multi-Anchor Projection Similarity (MAPS) metric. MAPS measures cross-modal similarity through the projection length of a target feature onto an anchor plane spanned by image and text features, thereby surpassing the constraints of conventional cosine similarity. Correspondingly, a MAPS contrastive loss is designed to enable joint representation learning in high-dimensional space. The proposed approach achieves state-of-the-art performance on vision-language geolocation benchmarks, significantly improving cross-modal retrieval accuracy.

cross-modal alignmentgeometric consistencyjoint image-text queries

Existing multimodal large language models struggle to dynamically switch between camera-, object-, or direction-centered reference frames during complex spatial reasoning, often leading to failures in multi-step inference. This work proposes Multiview Perspective Spatial Mapping (MPSM), which constructs query-aligned visual cognitive graphs and textual spatial graphs, and introduces a tool-guided egocentric reasoning strategy together with a cognitive graph distillation mechanism. For the first time, this approach enables consistent and efficient spatial reasoning without relying on external geometric pipelines. It achieves state-of-the-art performance on both single-image and multi-image benchmarks, and the distilled model substantially reduces dependence on external geometric processing while preserving strong reasoning capabilities.

egocentric reasoningmultimodal large language modelsreference frames

This work addresses the challenge of cross-view localization for robots operating in unseen urban environments using publicly available semantic maps such as OpenStreetMap. The proposed method leverages a vision-language model (VLM) to extract semantic landmarks from panoramic images and employs a lightweight matcher—distilled from the VLM via knowledge distillation—to efficiently align these observations with large-scale semantic maps. Temporal pose estimates are refined through Bayesian filtering. Trained solely on daytime data from a single city, the framework demonstrates robust generalization across eleven diverse environmental conditions, including nighttime and blizzard scenarios, and scales to map areas spanning hundreds of square kilometers. The authors also release a semantic dataset and source code to support reproducibility and further research.

city-scalecross-view localizationOpenStreetMap

Hot Scholars

MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
HV

Holger Voos

University of Luxembourg, SnT Automation & Robotics Research Group
Control EngineeringAutomationMobile Robotics
LC

Luca Carlone

Associate Professor, Massachusetts Institute of Technology
RoboticsRobot PerceptionComputer VisionEstimation and Inference
NN

Nassir Navab

Professor of Computer Science, Technische Universität München
GZ

Guangyao Zhai

Technical University of Munich; ETH Zurich
Generative AIEmbodied AI