3d representation learning

Designs and builds algorithms and models that extract, encode, and manipulate three‑dimensional structure and depth‑aware features—e.g., from voxels, point clouds, meshes, or RGB‑D inputs—to produce representations for 3D object detection, scene representation and reconstruction, generative 3D modeling, visualization, and explicit scene construction. Implements and analyzes operations for lifting and projecting features (3D-to-2D projection, depth-lifted object representations, feature projection layers), depth completion and depth supervision, voxelization and 3D feature extraction, and supporting traversal/search routines (including depth‑first search) used to train and evaluate learned 3D scene and object representations.

3drepresentationlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$180K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This paper presents a systematic review of learning-based 3D representations for tasks such as 3D reconstruction, novel view synthesis, and rendering, spanning from traditional explicit formulations—including meshes, point clouds, and voxels—to emerging implicit neural fields and primitive-based splatting techniques like 3D Gaussian Splatting. Emphasizing the evolutionary trajectory of 3D representations themselves, the work highlights the paradigm shift from explicit to implicit modeling, clarifying the mathematical formulations, strengths, limitations, and suitable application scenarios of each approach. Distinct from prior task-centric surveys, this study centers on representation as the core organizing principle, offering fresh insights for 3D/4D content generation and identifying key challenges and future research directions to serve as a theoretical reference for the computer graphics and vision communities.

3D representationsimplicit representationslearning-based methods

Recent Advance in 3D Object and Scene Generation: A Survey

Apr 16, 2025
XT
Xiang Tang
🏛️ Harbin Institute of Technology | Pengcheng Laboratory

Manual 3D modeling remains labor-intensive and time-consuming, failing to meet the rapidly growing demands of XR and metaverse applications. Method: This paper presents a systematic survey of state-of-the-art methods for static 3D object and scene generation, introducing— for the first time—a multidimensional classification and cross-comparative framework that jointly considers representation evolution (e.g., point clouds, meshes, NeRFs) and generative paradigms (e.g., supervised learning, diffusion models, 2D foundation model priors, procedural modeling). Contribution/Results: We identify key challenges—including geometric-semantic consistency and scalability—and establish a reusable, multi-axis evaluation framework. Our analysis clarifies technological trajectories and performance boundaries, providing both theoretical foundations and practical guidelines for industrial-grade 3D content generation, thereby advancing the paradigm shift in 3D content creation.

Addressing challenges in 3D content creation for XR/MetaverseOvercoming labor-intensive manual 3D modeling limitationsSurveying AI-driven 3D object and scene generation methods

Representation Learning for Point Cloud Understanding

Dec 05, 2025
SY
Siming Yan
🏛️ The University of Texas at Austin

This work addresses the limited representational capacity of point cloud learning by proposing a cross-modal representation learning framework that incorporates 2D visual priors. Methodologically, it innovatively leverages pre-trained 2D vision models to guide 3D point cloud network training via feature alignment—enabling effective 2D→3D knowledge transfer without naïve modality conversion. The framework integrates a point cloud encoder-decoder architecture, self-supervised pre-training, and primitive-level supervised segmentation into a unified learning paradigm. Extensive experiments on ScanNet and S3DIS benchmarks demonstrate substantial improvements in point cloud segmentation and scene understanding performance, validating the efficacy of 2D semantic priors in enhancing 3D representation learning. The proposed approach establishes a scalable, multimodal representation learning pathway for efficient 3D perception in autonomous driving and robotics applications.

Self-supervised learning methods for point cloud understandingSupervised learning for point cloud primitive segmentationTransfer learning from 2D to 3D using pre-trained models

The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.

3D visionbenchmark fragmentationdata representation

Towards Cross-device and Training-free Robotic Grasping in 3D Open World

Nov 27, 2024
WZ
Weiguang Zhao
🏛️ University of Liverpool | Duke Kunshan University | Xi’an-Jiaotong Liverpool University | Xi’an Jiaotong-Liverpool University

Existing approaches for robotic grasping in 3D open-world environments suffer from domain shift and poor generalization of clustering methods when handling cross-vendor cameras and robots. Method: We propose a training-free binary clustering framework that fuses multi-source, heterogeneous 3D point cloud segmentation outputs to achieve unsupervised clustering-based localization and robust grasping of unknown objects. Contribution/Results: Our work introduces the first training-free, plug-and-play paradigm for cross-device 3D point cloud processing, compatible with arbitrary 3D sensors. We design a lightweight binary clustering algorithm that eliminates reliance on prior distribution assumptions or scene-specific constraints. Evaluated across multiple robot platforms, diverse camera models, and cluttered, densely stacked scenes, our method achieves significant zero-shot grasping success rate improvements—demonstrating strong generalizability and deployment efficiency.

Address cross-device robotic grasping in 3D open worldExtend clustering methods to open-world settingsMinimize domain differences in point clouds from diverse cameras

Latest Papers

What's happening recently
View more

What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models

Dec 03, 2025
TD
Tianchen Deng
🏛️ Shanghai Jiao Tong University | University of Bonn | Nanyang Technological University | Chinese Academy of Sciences | University at Buffalo | I3A, University of Zaragoza

This paper presents a systematic survey of 3D scene representation methods for robotic tasks, addressing five core capabilities: perception, mapping, localization, navigation, and manipulation. It comparatively analyzes geometric and neural paradigms—including point clouds, voxels, signed distance fields (SDFs), neural radiance fields (NeRFs), and 3D Gaussian splatting—highlighting trade-offs in accuracy, computational efficiency, generalization, and semantic interpretability. Methodologically, it proposes a unified architectural pathway centered on 3D foundation models, integrating multimodal priors—particularly language and semantic knowledge—to enable high-level, embodied scene understanding. The work contributes an open-source evaluation benchmark and a modular implementation framework, offering the first structured taxonomy and evolutionary analysis spanning the full technical spectrum. This serves as both a theoretical reference and a practical development guide for advancing robotic 3D scene understanding.

Comparing geometric and neural representations like NeRF and foundation modelsEvaluating 3D scene representations for robotics across perception and navigation tasksExploring foundation models as a unified solution for future robotic applications

Existing 3D reconstruction methods struggle to balance generalization and practicality due to either inefficient per-scene optimization or reliance on category-specific training. This work proposes a feed-forward, output-representation-agnostic framework for 3D reconstruction, systematically addressing five core challenges: feature enhancement, geometry awareness, model efficiency, data augmentation, and temporal modeling. By unifying the analysis of image backbones, multi-view fusion mechanisms, and geometric priors—and integrating major datasets and evaluation benchmarks—it establishes a standardized benchmarking protocol. The study transcends differences in geometric representations, formulates a problem-driven, generalizable modeling paradigm, and outlines promising future directions in scalability, evaluation metrics, and world modeling.

3D scene representationcross-scene generalizationfeed-forward 3D reconstruction

This work addresses the challenges of weak spatial understanding and poor cross-view generalization in robotic systems operating under single-view constraints. To this end, the authors propose a unified representation–policy learning framework that introduces a novel single-view 3D pretraining paradigm, integrating point cloud reconstruction with feedforward Gaussian splatting. Through a multi-step knowledge distillation process, geometric-aware representations are effectively transferred to downstream manipulation policies. Evaluated on 12 RLBench tasks, the method achieves an average success rate surpassing the state of the art by 12.7%. Moreover, under large viewpoint shifts, it exhibits only a 29.7% drop in zero-shot success rate—significantly outperforming the state-of-the-art drop of 51.5%—demonstrating its strong viewpoint generalization capability.

3D visual representationsgeometric scene understandingsingle-view 3D reconstruction

This work proposes a fully automated method for reconstructing high-fidelity three-dimensional models from multiple orthogonal views of an object. The approach begins by extracting key control points using Harris corner detection, followed by generating mutually orthogonal bounding volumes through orthographic projection and constructing a 3D point cloud from their intersections. Subsequently, computational geometry algorithms are employed to recover the surface topology, and the resulting model is rendered using OpenGL for visualization. The entire pipeline operates without human intervention, achieving end-to-end reconstruction from 2D orthogonal projections to a complete 3D structure. The method demonstrates notable advantages in geometric consistency and reconstruction completeness compared to existing approaches.

3D modeling3D reconstructionautomated modeling

Advances in 4D Representation: Geometry, Motion, and Interaction

Oct 22, 2025
MZ
Mingrui Zhao
🏛️ Simon Fraser University | University of Alberta

Dynamic 3D scene modeling requires joint handling of geometric evolution, motion representation, and interactive semantics—yet existing 4D representations lack a unified conceptual framework addressing all three dimensions. Method: We propose the first taxonomy of 4D representations structured around the tripartite pillars of “geometry–motion–interaction.” We systematically survey mainstream approaches—including neural radiance fields, 3D Gaussian splatting, structured neural fields, and video foundation models—analyzing their limitations in generation and reconstruction tasks. We introduce a co-optimization framework for representation selection and task customization, emphasizing long-range motion modeling and multimodal data-driven design. We further investigate the integration potential and inherent boundaries of large language models and video foundation models in 4D understanding. Contribution/Results: Our work delivers a methodological taxonomy, a practical representation selection guide, dataset evaluation benchmarks, and an identification of critical research gaps—providing both theoretical foundations and actionable pathways for 4D generative AI.

Evaluating representations for geometry, motion, and interaction under different scenariosGuiding selection of appropriate 4D representations for specific tasksSurveying 4D generation and reconstruction methods for evolving 3D geometry

Hot Scholars

MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
GH

Gim Hee Lee

Associate Professor of Computer Science, National University of Singapore
Computer VisionRoboticsMachine Learning
ZC

Zhaopeng Cui

Zhejiang University
Computer VisionRoboticsComputer Graphics
JH

Junhui Hou

Department of Computer Science, City University of Hong Kong
Neural Spatial Computing
XH

Xiaoguang Han

Assistant Professor, The Chinese University of Hong Kong, Shenzhen
Computer VisionComputer Graphics