Score
Designs and builds algorithms and models that extract, encode, and manipulate three‑dimensional structure and depth‑aware features—e.g., from voxels, point clouds, meshes, or RGB‑D inputs—to produce representations for 3D object detection, scene representation and reconstruction, generative 3D modeling, visualization, and explicit scene construction. Implements and analyzes operations for lifting and projecting features (3D-to-2D projection, depth-lifted object representations, feature projection layers), depth completion and depth supervision, voxelization and 3D feature extraction, and supporting traversal/search routines (including depth‑first search) used to train and evaluate learned 3D scene and object representations.
This paper presents a systematic review of learning-based 3D representations for tasks such as 3D reconstruction, novel view synthesis, and rendering, spanning from traditional explicit formulations—including meshes, point clouds, and voxels—to emerging implicit neural fields and primitive-based splatting techniques like 3D Gaussian Splatting. Emphasizing the evolutionary trajectory of 3D representations themselves, the work highlights the paradigm shift from explicit to implicit modeling, clarifying the mathematical formulations, strengths, limitations, and suitable application scenarios of each approach. Distinct from prior task-centric surveys, this study centers on representation as the core organizing principle, offering fresh insights for 3D/4D content generation and identifying key challenges and future research directions to serve as a theoretical reference for the computer graphics and vision communities.
Manual 3D modeling remains labor-intensive and time-consuming, failing to meet the rapidly growing demands of XR and metaverse applications. Method: This paper presents a systematic survey of state-of-the-art methods for static 3D object and scene generation, introducing— for the first time—a multidimensional classification and cross-comparative framework that jointly considers representation evolution (e.g., point clouds, meshes, NeRFs) and generative paradigms (e.g., supervised learning, diffusion models, 2D foundation model priors, procedural modeling). Contribution/Results: We identify key challenges—including geometric-semantic consistency and scalability—and establish a reusable, multi-axis evaluation framework. Our analysis clarifies technological trajectories and performance boundaries, providing both theoretical foundations and practical guidelines for industrial-grade 3D content generation, thereby advancing the paradigm shift in 3D content creation.
This work addresses the limited representational capacity of point cloud learning by proposing a cross-modal representation learning framework that incorporates 2D visual priors. Methodologically, it innovatively leverages pre-trained 2D vision models to guide 3D point cloud network training via feature alignment—enabling effective 2D→3D knowledge transfer without naïve modality conversion. The framework integrates a point cloud encoder-decoder architecture, self-supervised pre-training, and primitive-level supervised segmentation into a unified learning paradigm. Extensive experiments on ScanNet and S3DIS benchmarks demonstrate substantial improvements in point cloud segmentation and scene understanding performance, validating the efficacy of 2D semantic priors in enhancing 3D representation learning. The proposed approach establishes a scalable, multimodal representation learning pathway for efficient 3D perception in autonomous driving and robotics applications.
The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.
Existing approaches for robotic grasping in 3D open-world environments suffer from domain shift and poor generalization of clustering methods when handling cross-vendor cameras and robots. Method: We propose a training-free binary clustering framework that fuses multi-source, heterogeneous 3D point cloud segmentation outputs to achieve unsupervised clustering-based localization and robust grasping of unknown objects. Contribution/Results: Our work introduces the first training-free, plug-and-play paradigm for cross-device 3D point cloud processing, compatible with arbitrary 3D sensors. We design a lightweight binary clustering algorithm that eliminates reliance on prior distribution assumptions or scene-specific constraints. Evaluated across multiple robot platforms, diverse camera models, and cluttered, densely stacked scenes, our method achieves significant zero-shot grasping success rate improvements—demonstrating strong generalizability and deployment efficiency.
This paper presents a systematic survey of 3D scene representation methods for robotic tasks, addressing five core capabilities: perception, mapping, localization, navigation, and manipulation. It comparatively analyzes geometric and neural paradigms—including point clouds, voxels, signed distance fields (SDFs), neural radiance fields (NeRFs), and 3D Gaussian splatting—highlighting trade-offs in accuracy, computational efficiency, generalization, and semantic interpretability. Methodologically, it proposes a unified architectural pathway centered on 3D foundation models, integrating multimodal priors—particularly language and semantic knowledge—to enable high-level, embodied scene understanding. The work contributes an open-source evaluation benchmark and a modular implementation framework, offering the first structured taxonomy and evolutionary analysis spanning the full technical spectrum. This serves as both a theoretical reference and a practical development guide for advancing robotic 3D scene understanding.
Existing 3D reconstruction methods struggle to balance generalization and practicality due to either inefficient per-scene optimization or reliance on category-specific training. This work proposes a feed-forward, output-representation-agnostic framework for 3D reconstruction, systematically addressing five core challenges: feature enhancement, geometry awareness, model efficiency, data augmentation, and temporal modeling. By unifying the analysis of image backbones, multi-view fusion mechanisms, and geometric priors—and integrating major datasets and evaluation benchmarks—it establishes a standardized benchmarking protocol. The study transcends differences in geometric representations, formulates a problem-driven, generalizable modeling paradigm, and outlines promising future directions in scalability, evaluation metrics, and world modeling.
This work addresses the challenges of weak spatial understanding and poor cross-view generalization in robotic systems operating under single-view constraints. To this end, the authors propose a unified representation–policy learning framework that introduces a novel single-view 3D pretraining paradigm, integrating point cloud reconstruction with feedforward Gaussian splatting. Through a multi-step knowledge distillation process, geometric-aware representations are effectively transferred to downstream manipulation policies. Evaluated on 12 RLBench tasks, the method achieves an average success rate surpassing the state of the art by 12.7%. Moreover, under large viewpoint shifts, it exhibits only a 29.7% drop in zero-shot success rate—significantly outperforming the state-of-the-art drop of 51.5%—demonstrating its strong viewpoint generalization capability.
This work proposes a fully automated method for reconstructing high-fidelity three-dimensional models from multiple orthogonal views of an object. The approach begins by extracting key control points using Harris corner detection, followed by generating mutually orthogonal bounding volumes through orthographic projection and constructing a 3D point cloud from their intersections. Subsequently, computational geometry algorithms are employed to recover the surface topology, and the resulting model is rendered using OpenGL for visualization. The entire pipeline operates without human intervention, achieving end-to-end reconstruction from 2D orthogonal projections to a complete 3D structure. The method demonstrates notable advantages in geometric consistency and reconstruction completeness compared to existing approaches.
Dynamic 3D scene modeling requires joint handling of geometric evolution, motion representation, and interactive semantics—yet existing 4D representations lack a unified conceptual framework addressing all three dimensions. Method: We propose the first taxonomy of 4D representations structured around the tripartite pillars of “geometry–motion–interaction.” We systematically survey mainstream approaches—including neural radiance fields, 3D Gaussian splatting, structured neural fields, and video foundation models—analyzing their limitations in generation and reconstruction tasks. We introduce a co-optimization framework for representation selection and task customization, emphasizing long-range motion modeling and multimodal data-driven design. We further investigate the integration potential and inherent boundaries of large language models and video foundation models in 4D understanding. Contribution/Results: Our work delivers a methodological taxonomy, a practical representation selection guide, dataset evaluation benchmarks, and an identification of critical research gaps—providing both theoretical foundations and actionable pathways for 4D generative AI.