Score
Designs and implements methods and pipelines that recover the spatial structure, appearance, semantics, and optionally dynamics of an environment from observations (for example images, depth, lidar, or other sensor streams), producing explicit representations such as 3D meshes, point clouds, occupancy/volumetric grids, semantic maps, or scene graphs. Also develops evaluation and analysis techniques to measure reconstruction accuracy, completeness, consistency across modalities, and temporal alignment.
This paper addresses the lack of a hierarchical analytical framework in existing surveys on 4D scene reconstruction. We propose the first five-level layered taxonomy that spans the full cognitive evolution path: from low-level 3D geometric reconstruction, through motion modeling and interaction reasoning, to physical law learning and causal inference. Methodologically, we unify multi-view geometry, temporal neural networks, 3D representation learning, and physics-based simulation to systematically integrate diverse paradigms for dynamic scene modeling. Our contributions are threefold: (1) introducing the first structurally coherent, semantically progressive classification framework for 4D spatial intelligence, filling a critical gap in hierarchical survey literature; (2) distilling core challenges and developmental trajectories at each level; and (3) establishing an open-source project page to continuously track advances, thereby fostering systematic knowledge organization and community-wide sharing.
Existing world modeling research predominantly focuses on 2D image/video generation, neglecting large-scale scene modeling using native 3D/4D representations—such as RGB-D, occupancy grids, and LiDAR point clouds—and lacks a unified definition and systematic taxonomy. Method: This paper introduces, for the first time, a standardized definition and a structured classification framework for 3D/4D world models, systematically categorizing generative paradigms into VideoGen, OccGen, and LiDARGen. It integrates generative modeling, 3D perception, and spatiotemporal modeling techniques, and synthesizes evaluation metrics and benchmark datasets. Contribution/Results: As the first comprehensive survey in this emerging field, it establishes WorldBench—an open-source literature platform—thereby filling a critical theoretical gap and providing foundational guidance and standardization pathways for 3D/4D world modeling research.
This paper presents a systematic review of learning-based 3D representations for tasks such as 3D reconstruction, novel view synthesis, and rendering, spanning from traditional explicit formulations—including meshes, point clouds, and voxels—to emerging implicit neural fields and primitive-based splatting techniques like 3D Gaussian Splatting. Emphasizing the evolutionary trajectory of 3D representations themselves, the work highlights the paradigm shift from explicit to implicit modeling, clarifying the mathematical formulations, strengths, limitations, and suitable application scenarios of each approach. Distinct from prior task-centric surveys, this study centers on representation as the core organizing principle, offering fresh insights for 3D/4D content generation and identifying key challenges and future research directions to serve as a theoretical reference for the computer graphics and vision communities.
To address the challenge of balancing geometric accuracy and visual fidelity in large-scale 3D spatial data for defense applications, this paper proposes a hierarchical hybrid representation architecture integrating classical geometric modeling with neural rendering. The architecture employs triangle meshes and voxel grids to ensure foundational geometric fidelity, while leveraging 3D Gaussian splatting and Neural Radiance Fields (NeRF) for high-fidelity photorealistic rendering at the upper layer. A unified scene management framework enables multi-granularity co-optimization across representations. Compared to purely geometric or purely neural approaches, our method achieves significant improvements: 2.1× acceleration in computational efficiency and +3.7 dB PSNR gain in rendering quality—particularly beneficial for line-of-sight analysis, physics-based simulation, and real-time visualization. It supports scalable modeling and interactive rendering of scenes with up to hundreds of millions of polygons, establishing a new paradigm for military digital twins that simultaneously delivers geometric precision, computational efficiency, and visual realism.
This work addresses the challenge of degraded 3D surface reconstruction accuracy caused by missing geometric information in LiDAR point clouds due to limited scanning range and occlusions. To tackle this issue, the authors propose a reconstruction method based on plane classification and priority-driven growth. The approach categorizes scene planes into three visibility classes—highly visible, partially visible, and invisible—and employs a hierarchical spatial partitioning scheme. Coupled with a min-cut optimization strategy, it generates compact, watertight polygonal models that effectively recover missing geometric details. Evaluated on public datasets, the method significantly outperforms current state-of-the-art techniques, achieving higher reconstruction fidelity while preserving model compactness.
The absence of publicly available, semantically annotated point clouds generated via Structure-from-Motion (SfM) for complex scenes—such as forests—hinders the training and evaluation of deep learning models. Method: We propose an end-to-end SfM semantic point cloud generation pipeline: (i) synthesizing RGB images with pixel-level semantic masks using a custom forest simulator; (ii) modifying COLMAP to enable cross-view semantic fidelity preservation during SfM reconstruction—the first such adaptation; and (iii) incorporating multi-view geometric constraints and joint semantic-geometric optimization to enhance semantic consistency under dense forest structures. Contribution/Results: We release the first publicly available, reproducible SfM-based semantic point cloud dataset, annotated with categories including trunks, canopies, and ground. Our method improves semantic projection accuracy by 37% and significantly boosts the generalization performance of downstream semantic segmentation models on real-world SfM point clouds.
To address insufficient geometric-semantic joint understanding of robots in unstructured environments, this paper proposes an end-to-end modular RGB-D understanding pipeline. Our method introduces a novel hybrid mask generation mechanism integrating SAM2 with a lightweight semantic classifier for pixel-level semantic segmentation and instance awareness; incorporates ReID-enhanced cross-frame human tracking and semantic-weighted TSDF point cloud fusion to ensure geometric fidelity and semantic consistency; and outputs structured scene representations in USD format. Evaluated on ADE20K, our approach achieves 47.0% mIoU—surpassing SegFormer and OneFormer in boundary accuracy—and attains a reconstruction error of only 25.3 mm. It runs 1.81× faster than prior methods and has been validated for deployment feasibility on real-world Kinect data.
The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.
This work addresses the challenges of sparse LiDAR point clouds, geometric drift, and fixed fusion parameters in large-scale indoor scenes—issues that commonly lead to mesh holes, over-smoothing, and boundary artifacts. To overcome these limitations, the authors propose a modular, incremental RGB-LiDAR fusion framework that, for the first time, integrates per-frame semantic labels generated by vision foundation models into truncated signed distance function (TSDF) voxels in an incremental manner, enabling semantic-aware high-fidelity mesh reconstruction. The system cohesively combines visual semantic labeling, LiDAR-inertial odometry mapping, semantic-aware TSDF fusion, and Marching Cubes surface extraction. Evaluated on the Oxford Spires dataset, the method achieves superior geometric reconstruction accuracy compared to state-of-the-art approaches such as ImMesh and Voxblox, and produces semantic meshes directly suitable for Universal Scene Description (USD) asset creation and extended reality (XR) applications.
This work addresses the insufficient integration of learning-based methods and geometric constraints in camera pose and scene structure estimation by proposing a modular framework. The approach first employs a learning model (VGGT) to generate initial hypotheses for depth and relative pose, which are subsequently refined and validated using classical geometric algorithms such as point-to-plane RGB-D ICP. Crucially, the framework explicitly distinguishes the roles of learning as a “proposer” and geometry as a “referee,” emphasizing that the geometric module serves not merely as post-processing but as an essential mechanism for verifying and integrating learned outputs. Experiments on the TUM RGB-D dataset demonstrate that, in moderately challenging rigid scenes, the system significantly outperforms both purely learning-based and purely geometric baselines when the learned depth aligns geometrically with the camera intrinsics and undergoes optimization by the geometric backend.
Existing 3D scene reconstruction methods struggle to meet the demands of interactive graphics applications for high-quality, cross-view consistent semantic segmentation. This work proposes a novel approach based on Radiant Foam—a voxelized Voronoi grid—by introducing an explicit semantic feature field at the cell level and, for the first time, directly incorporating spatial regularization into this field to achieve joint spatial-semantic decomposition. This design significantly enhances cross-view consistency and effectively mitigates artifacts caused by occlusions and inconsistent supervision. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches such as Gaussian Grouping and SAGA in both object-level semantic segmentation accuracy and consistency.