Score
Designs, implements, and analyzes pipelines that voxelize 3D sensor data (point clouds, depth frames) and construct voxel backbones that project multi-view observations into a shared voxel space using depth and camera pose. These systems perform probabilistic fusion and source-aware semantic labeling to produce persistent volumetric semantic maps that support incremental queries, redundancy detection, voxel-based alignment, multi-resolution/local refinement, and other voxel-grid processing operations.
Existing world modeling research predominantly focuses on 2D image/video generation, neglecting large-scale scene modeling using native 3D/4D representations—such as RGB-D, occupancy grids, and LiDAR point clouds—and lacks a unified definition and systematic taxonomy. Method: This paper introduces, for the first time, a standardized definition and a structured classification framework for 3D/4D world models, systematically categorizing generative paradigms into VideoGen, OccGen, and LiDARGen. It integrates generative modeling, 3D perception, and spatiotemporal modeling techniques, and synthesizes evaluation metrics and benchmark datasets. Contribution/Results: As the first comprehensive survey in this emerging field, it establishes WorldBench—an open-source literature platform—thereby filling a critical theoretical gap and providing foundational guidance and standardization pathways for 3D/4D world modeling research.
This work addresses the lack of environment representation methods for mobile manipulators under edge computing constraints that simultaneously ensure timeliness, semantic richness, and geometric reliability. The authors propose KRVF, a task-oriented voxelized world representation that integrates occupancy, color, semantic evidence, temporal recency, and data provenance to enable closed-loop feedback between mapping and perception. Key innovations include decoupling observed occupancy from semantic priors to enhance object reasoning robustness under depth failure, introducing a map-prior-driven depth inpainting mechanism, and designing a task-aware semantic query interface. Built upon semantic voxel representations, source-aware fusion, and real-time ROS 2 integration, the system efficiently supports semantic object retrieval and grasp candidate generation on edge devices.
To address insufficient geometric-semantic joint understanding of robots in unstructured environments, this paper proposes an end-to-end modular RGB-D understanding pipeline. Our method introduces a novel hybrid mask generation mechanism integrating SAM2 with a lightweight semantic classifier for pixel-level semantic segmentation and instance awareness; incorporates ReID-enhanced cross-frame human tracking and semantic-weighted TSDF point cloud fusion to ensure geometric fidelity and semantic consistency; and outputs structured scene representations in USD format. Evaluated on ADE20K, our approach achieves 47.0% mIoU—surpassing SegFormer and OneFormer in boundary accuracy—and attains a reconstruction error of only 25.3 mm. It runs 1.81× faster than prior methods and has been validated for deployment feasibility on real-world Kinect data.
This work addresses the challenge of achieving open-vocabulary 3D semantic mapping in large-scale environments, where existing methods often fall short. The authors propose an online two-layer semantic mapping framework that unifies voxel-level dense semantics and instance-level open-vocabulary representations within a shared voxel map for the first time. By introducing a cross-layer semantic fusion mechanism within a sliding window, the method jointly optimizes the quality of both semantic layers. Coupled with multi-view semantic embeddings and voxelized 3D mapping, this approach significantly enhances accuracy and scalability. The framework demonstrates high-fidelity, generalizable open-vocabulary semantic mapping on standard 3D semantic segmentation benchmarks as well as large, multi-floor real-world scenes, effectively extending to unseen semantic concepts.
This work addresses the limitation of existing open-vocabulary 3D scene graph methods, which rely on a sequential “reconstruct-then-enrich” pipeline that precludes real-time querying during exploration. To overcome this, the authors propose an asynchronous architecture that decouples lightweight online mapping from heavyweight semantic refinement, executing them in parallel. A probabilistic voxel backbone maintains object identity consistency, while a background visual-language model (VLM) agent incrementally enriches semantic annotations. The approach further incorporates semantic loop closure to eliminate redundant trajectories and introduces a multi-objective frame scheduler to reduce VLM computational overhead. This is the first method to enable queryable scene graphs during active exploration, achieving state-of-the-art semantic segmentation performance on ScanNet and Replica, and significantly outperforming prior art by 15.3–18.8 A@0.25 on the Sr3D+, Nr3D, and ScanRefer benchmarks.
In LiDAR-camera fusion 3D detection, the sparsity of voxel features and the density of image features hinder cross-modal alignment and lead to semantic and continuity information loss—particularly for distant objects. To address this, we propose a voxel-space image feature processing framework: (1) image features are projected onto the voxel grid via voxelization; (2) pixel-level and patch-level multi-granularity features are extracted; (3) a cross-modal self-attention fusion mechanism enables fine-grained alignment of image features within the voxel space; and (4) a foreground-aware importance weighting module suppresses background interference. Evaluated on KITTI, our method achieves a +3.20% improvement in AP@0.7 for hard-category car detection over the Voxel R-CNN baseline. Furthermore, it demonstrates strong generalizability on nuScenes, validating its robustness across diverse driving scenarios and sensor configurations.
本文提出VoxelFix方法,通过基于图的模型直接从已完成的3D体素地图中纠正语义标签,提高地图准确性,用于解决自动构建语义3D地图中的错误问题。
The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.
该研究提出一种结合外部校准相机、对象检测、持续跟踪及本体驱动语义更新的混合管道,以构建动态语义世界模型,解决机器人在复杂环境中的交互问题。
该研究解决了3D场景图中不确定性表示和传播问题,通过引入概率场景图(PSG)及高斯层次图(HGG),实现了实时感知与精确定位。
This work investigates how prior information from a fixed external RGB camera can enhance the initial completeness and exploration efficiency of 3D scene graphs in robotic active exploration. The method models external camera observations as a Common Prior Map (CPM), constructing semantic and geometric priors before exploration begins. A hardware-agnostic, RGB-only multi-view fusion framework seamlessly integrates both onboard and external viewpoints to enable incremental 3D scene graph generation. The key innovation lies in leveraging a fixed external camera as a universal prior source without requiring hardware modifications, coupled with a semantic uncertainty–guided active exploration strategy driven by partial scene graphs. Experiments demonstrate that a single external camera can improve initial object recall by up to 79%, substantially boosting both exploration efficiency and scene graph completeness.