Score
Designs and implements probabilistic 3D scene-graph representations and algorithms that fuse incremental 3D sensor observations into a graph of entities and structural relations, maintaining posterior beliefs and per-node uncertainty for nodes and edges. Builds systems to construct, update, and reason over uncertainty-aware semantic scene graphs (including integration of priors such as 2D occupancy maps) to support incremental semantic mapping and decision-making.
该研究解决了3D场景图中不确定性表示和传播问题,通过引入概率场景图(PSG)及高斯层次图(HGG),实现了实时感知与精确定位。
Existing 3D Semantic Scene Graph (3DSSG) methods rely on complete scene reconstruction and single-sensor input, limiting their applicability to real-world incremental and dynamic modeling scenarios. To address this, we propose an end-to-end incremental 3DSSG prediction framework. Our approach introduces a heterogeneous graph neural network that directly incorporates historical observations into the message-passing process, enabling joint global–local representation learning. It fuses multimodal inputs—including RGB-D data and textual prompts—and leverages CLIP-based semantic embeddings for cross-modal alignment. Crucially, the method operates without requiring full-scene reconstruction, thereby significantly improving generalization and scalability in partially observable, continuously evolving environments. We validate its effectiveness on standard 3DSSG benchmarks, demonstrating robust performance under incremental observation settings. This work establishes a deployable foundation for long-horizon intelligent interaction through incremental semantic understanding.
Existing online 3D scene graph generation methods overlook uncertainties inherent in observations, 2D models, and 3D representations, leading to overly deterministic fusion processes. This work proposes a plug-and-play, training-free, uncertainty-aware fusion framework that, for the first time, integrates explicit uncertainty modeling into online 3D scene graph construction. The approach jointly models semantic and spatial factors through probabilistic likelihoods to infer node associations, accumulates categorical and relational evidence using Dirichlet-based evidential reasoning, and replaces conventional binary gating with probabilistic association. Compatible with diverse 3D representations—including Gaussian and voxel-based formats—the method achieves state-of-the-art performance on both 3DSSG and ReplicaSSG benchmarks while maintaining real-time inference speed.
This work proposes a novel approach to semantic mapping by adopting 3D Semantic Scene Graphs (3DSSGs) as the core representational layer, addressing the inconsistency and limited scalability of existing methods that decouple perception from semantic representation in large-scale real-world environments. By incrementally constructing and updating the graph structure in real time during exploration, the method bridges the gap between raw sensor data and high-level knowledge systems. It leverages incremental scene graph prediction, spatially anchored explicit graph representations, and a unified mechanism for integrating flat and hierarchical topologies. This framework seamlessly incorporates external knowledge sources—such as knowledge graphs, ontologies, and large language models—enabling efficient, consistent, and interpretable open-set semantic maps over long-term, large-scale operation, thereby significantly enhancing agent trustworthiness and alignment with human conceptual understanding in complex environments.
This work addresses the limitation of existing 3D scene graph methods, which treat perception as a post-processing step on static datasets and decouple scene understanding from observation planning, thereby hindering long-term, incremental environment modeling for robots. To overcome this, the paper introduces an online semantic exploration framework that, for the first time, formulates semantic scene completeness as an active optimization objective. The approach employs an uncertainty-guided traversal strategy to dynamically balance semantic verification, geometric coverage, and motion cost. By fusing RGB-D observations with a prior 2D occupancy map, it incrementally constructs an uncertainty-aware 3D scene graph encoding open-vocabulary object label posteriors and structural relational edges, which in turn drives closed-loop path planning. The system autonomously revisits semantically ambiguous regions and explores unknown spaces, enabling continuous, human-intervention-free patrolling, updating, and reasoning.
This work addresses the fragmented state of 3D scene graph research, hindered by the absence of a unified definition, construction pipeline, and evaluation protocol, which impedes method comparison and real-world deployment. The paper presents the first formal definition and a cohesive theoretical framework for 3D scene graphs, systematically analyzing key modeling choices—including node and edge attributes, hierarchical structure, dynamic modeling, and functional awareness—and clarifying the terminology and mainstream methodologies that map perceptual data to scene graphs. Through a comprehensive literature review, taxonomic analysis, and technical comparison, it delineates the core components and evolutionary trajectories of the field, identifies critical research gaps in geometry–semantics integration, relational reasoning, dynamic modeling, and task-driven evaluation, and introduces an accompanying knowledge website to provide a clear roadmap for algorithm development, benchmarking, and standardized collaboration.
This work addresses the limitations of existing 3D scene graph methods, which are constrained by predefined relationship categories and struggle to capture open-ended semantics and causal connections. To overcome this, the authors propose a novel framework that integrates vision-language models (VLMs) with large language models (LLMs) to construct a hierarchical forest of 3D semantic scene graphs. The VLM extracts instance-level nodes and geometry-aware relationships, while the LLM performs high-level reasoning to generate abstract concepts and open-vocabulary semantic associations. This approach transcends the confines of closed relationship sets, substantially enhancing the semantic depth and expressiveness of scene representations. Experiments on uHumans2 and ScanNet demonstrate improved accuracy in relationship generation, and real-world deployment on a Spot robot successfully enables open-vocabulary object retrieval in physical environments.
This study addresses the lack of structured reasoning in semantic mapping via 3D Gaussian Splatting (3DGS) and the difficulty of leveraging dense semantics for scene graph construction. We propose a unified framework that directly constructs persistent 3D scene graphs from online Gaussian semantic maps. A core innovation lies in tightly coupling dense semantics with object-centric representations by introducing a reliability-aware semantic field, which enables confidence-guided object extraction and supports incremental graph updates. The proposed method demonstrates superior performance on standard benchmarks and real-world robotic experiments, effectively balancing geometric fidelity with efficient open-vocabulary perception. Furthermore, it significantly enhances downstream tasks such as language-guided localization.
本文提出一种框架,利用3D场景图中存储的检测置信度和嵌入信息来估计语义不确定性,并通过层次结构传播,从而提高对象检索准确性并降低房间级别断言的错误。
研究通过分析五个户外机器人数据集,使用Terra 3DSG作为案例,探讨了开放环境下3D场景图在语义点嵌入、导航等方面的挑战,并提出新的度量标准评估一致性。