Score
Designs and implements hierarchical spatial occupancy representations that combine local submaps (e.g., room-level) and a top-level building graph, storing occupancy probabilities and semantic labels at multiple scales. Work includes organizing memory into local caches, aggregating and fusing submaps into the building-level graph, and converting continuous sensor estimates to discrete occupancy (e.g., gaussian-to-occupancy splatting) for efficient querying and updates.
To address the high memory overhead and computational inefficiency in semantic mapping for embodied agents operating long-term in unknown indoor environments, this paper presents the first comprehensive, multi-dimensional survey tailored to indoor scenarios. We systematically classify existing approaches along two orthogonal dimensions—structural representation (grid-based, topological, point-cloud, or hybrid) and information type (implicit features vs. explicit data)—and critically analyze state-of-the-art methods, including deep semantic segmentation, SLAM-semantic integration, cross-modal alignment, graph neural network (GNN)-based modeling, and 3D reconstruction, identifying their technical limits and bottlenecks. We propose an evolutionary paradigm for semantic maps characterized by open-vocabulary support, queryability, and task-agnosticism. Finally, we distill four key future research directions. This work establishes a theoretical framework and technical roadmap for advancing semantic understanding, long-horizon navigation, and task planning in embodied intelligence.
This work addresses the challenge of achieving long-term consistent semantic occupancy mapping across interconnected indoor spaces. To this end, we propose GEM-Occ, a novel framework that introduces a Gaussian evidence memory–based mechanism for semantic occupancy mapping. Our approach converts local visual predictions into semantic Gaussian occupancy and free-space ray evidence, which are then integrated into a hierarchical persistent memory structure—comprising local buffers, room-level subgraphs, and building-scale graphs—via a visibility- and uncertainty-aware causal update strategy. The framework further enables efficient Gaussian-to-occupancy splatting for rapid querying. Evaluated on the HIOcc benchmark, GEM-Occ substantially outperforms existing methods in local accuracy, online stability, free-space reasoning, revisit consistency, and building-scale scalability.
Zero-shot semantic 3D occupancy reconstruction from unlabeled multi-sensor data remains challenging due to the absence of semantic supervision and effective joint geometric-semantic modeling. Method: This paper proposes a scene-aware geometry-semantic joint reconstruction framework. Its core components include: (1) a novel Gaussian lattice modeling paradigm that jointly perceives semantics and geometry; (2) a cumulative Gaussian voxel splatting algorithm enabling efficient, differentiable voxelization; and (3) integration of vision-language model priors, LiDAR-guided geometric constraints, and jointly parameterized Gaussian representations—entirely without semantic annotations. Results: Under zero-shot settings, our method achieves state-of-the-art performance, substantially outperforming existing self-supervised approaches while attaining semantic occupancy accuracy close to fully supervised baselines. It establishes a new paradigm for open-world 3D understanding.
本文提出了一种基于采样的方法,通过统计估计而非穷举递归构建信息驱动的多分辨率概率占据网格层次表示,解决了大规模网格计算复杂性问题。
Existing LiDAR occupancy grid prediction methods predominantly rely on deterministic, grid-level optimization, often yielding physically implausible and scene-inconsistent artifacts that compromise safety-critical autonomous navigation. To address this, we propose the first latent-space disentangled generative occupancy prediction framework: it decouples representation learning from stochastic prediction, explicitly modeling uncertainty in the latent space while supporting multimodal conditional inputs—including RGB images and high-definition maps. Our approach employs a hybrid VAE-GAN architecture, enabling end-to-end training and zero-shot cross-platform transfer. Evaluated on NuScenes, Waymo Open Dataset, and a proprietary real-world vehicle dataset, our method achieves state-of-the-art performance, significantly improving prediction fidelity, physical plausibility, and scene consistency—key requirements for robust autonomous driving systems.
To address insufficient environmental understanding in indoor intelligent robot navigation, this paper proposes a hierarchical 3D scene graph (3DSG) construction framework. The 3DSG comprises three layers: a metric-semantic foundation layer, an object-level point cloud/visual representation layer, and high-level semantic nodes (e.g., rooms, floors). Innovatively, it introduces large language models (LLMs) for the first time to automate labeling across all hierarchical nodes—including room-level and above—and designs an LLM-based voting mechanism for room classification, significantly improving accuracy and robustness of high-level semantic annotation. The method integrates multi-modal RGB-D and point cloud perception, geometry-semantic joint modeling, prompt-engineered LLM reasoning, and hierarchical graph structure optimization. Experiments demonstrate that the generated 3DSG exhibits rich semantics and structural completeness, achieving substantial improvements over baselines in geometric-semantic fusion accuracy, context-aware navigation, and task planning capability.
This work addresses the challenges of trajectory drift accumulation and computationally expensive global consistency optimization in large-scale SLAM, as well as limitations of existing discrete grid-based submap stitching methods—such as discontinuous gradients and neglect of occupancy uncertainty—by introducing the first continuous probabilistic submap stitching framework. The method jointly optimizes submap poses and a global occupancy field in an implicit log-odds space, compressing raw observations into informative sufficient statistics via sparse Bayesian inference and incorporating a variance-weighting mechanism to preserve posterior uncertainty. It enables analytical Jacobian computation and directly yields an optimal global map with closed-form mean and variance upon pose convergence. Experiments demonstrate significant improvements over state-of-the-art approaches in both simulated and real large-scale environments, achieving higher pose accuracy, enhanced global consistency, greater map compactness, and better-calibrated uncertainty.
Existing high-resolution 3D occupancy prediction methods are hindered by substantial computational costs and memory bottlenecks, making it challenging to achieve both accuracy and real-time performance. This work proposes a progressive multi-scale Gaussian occupancy prediction framework that overcomes the memory limitations of dense voxel representations through a hierarchical Gaussian seeding mechanism, which constructs a coarse-to-fine structure of Gaussian primitives. By integrating hierarchical Gaussian modeling, multi-scale feature fusion, and a hybrid sparse-dense representation, our approach enables end-to-end trainable and highly efficient inference. To support this research, we introduce TJScenes, the first panoramic six-camera dataset with fine-grained 0.1-meter occupancy annotations. Evaluated on both Occ3D-nuScenes and TJScenes, our method achieves state-of-the-art geometry reconstruction accuracy while significantly reducing inference latency, establishing a new Pareto frontier in the trade-off between efficiency and precision.
This work addresses the lack of a consistent geometric evaluation criterion for room nodes in existing hierarchical 3D scene graphs, which leads to structural inconsistencies. To resolve this, the authors propose a novel approach grounded in occupancy decomposition that tracks free-space regions and generates polygonal footprints, thereby introducing free space as the first geometric anchor for room nodes. This strategy unifies both the construction and evaluation of the room layer. By integrating occupancy decomposition, free-space tracking, polygon generation, and matching with Matterport3D room instances, the method significantly improves room recall across twelve scenes. Although precision experiences a slight decline, the results highlight wall alignment boundaries as a persistent challenge across diverse environments.
This work addresses the challenge that existing 3D semantic occupancy prediction methods struggle to handle heterogeneous indoor and outdoor scenes within a unified framework. To this end, we introduce the cross-scene 3D semantic occupancy prediction task and propose OccAnyScene, a novel framework built upon a pretrained foundation model. OccAnyScene leverages pixel-aligned frustum feature aggregation and a frustum-parameterized Gaussian decoding mechanism to adaptively reconstruct scenes with varying camera configurations and spatial scales. Our approach is the first to enable a single model to consistently process diverse environments while preserving both metric consistency and scene adaptability. It achieves state-of-the-art performance with mIoU scores of 59.92% on Occ-ScanNet (indoor) and 23.06% on SurroundOcc-nuScenes (outdoor), setting new benchmarks for cross-scene 3D semantic occupancy prediction.
该研究解决了3D场景图中不确定性表示和传播问题,通过引入概率场景图(PSG)及高斯层次图(HGG),实现了实时感知与精确定位。