2d-to-3d lifting

Design and implement algorithms and modules that convert 2D image- or object-level representations into 3D geometries or volumetric/feature grids (point clouds, voxels, meshes, or lifted latent volumes), including unprojection, alignment, and encoding strategies for single-view or multi-view input. Build and analyze pipelines that lift image features into 3D latent spaces and outputs (single-view mesh/voxel recovery, point-cloud/mesh reconstruction), integrate pretrained 2D encoders with 3D decoders, condition synthesis on viewpoint, and incorporate guidance such as diffusion priors or unsupervised pose lifting to produce and align 3D representations.

2d-to-3dlifting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey

Jul 19, 2025
JZ
Jiahui Zhang
🏛️ Nanyang Technological University | California Institute of Technology | Westlake University | University of Oxford | Nanjing University | University of Cambridge | Hillbot | University of California, San Diego | Max Planck Institute for Informatics | Harvard University | Massachusetts Institute of Technology

Motivated by the urgent demands of AR/VR and digital twin applications for fast, generalizable, and deployment-friendly 3D reconstruction and novel view synthesis, this paper presents a systematic survey of feedforward deep learning methods—covering dominant representations including point clouds, 3D Gaussian splatting, and neural radiance fields—and focuses on three key challenges: pose-free input, dynamic scene modeling, and 3D-aware content generation. We propose the first unified taxonomy tailored to the feedforward paradigm, revealing inherent trade-offs between inference efficiency and cross-scene generalization. By integrating self-supervised learning, differentiable rendering, and multimodal input strategies—and leveraging standardized evaluation protocols and large-scale benchmarks—we comprehensively assess accuracy, latency, and robustness. Our analysis provides principled guidance and empirically grounded technology selection criteria for industrial-grade 3D vision systems.

Addressing limitations of traditional iterative optimization methodsExploring applications in AR, VR, and digital humansSurveying feed-forward 3D reconstruction and view synthesis techniques

Must-Read Papers

Most classic and influential ideas
View more

This paper presents a systematic review of learning-based 3D representations for tasks such as 3D reconstruction, novel view synthesis, and rendering, spanning from traditional explicit formulations—including meshes, point clouds, and voxels—to emerging implicit neural fields and primitive-based splatting techniques like 3D Gaussian Splatting. Emphasizing the evolutionary trajectory of 3D representations themselves, the work highlights the paradigm shift from explicit to implicit modeling, clarifying the mathematical formulations, strengths, limitations, and suitable application scenarios of each approach. Distinct from prior task-centric surveys, this study centers on representation as the core organizing principle, offering fresh insights for 3D/4D content generation and identifying key challenges and future research directions to serve as a theoretical reference for the computer graphics and vision communities.

3D representationsimplicit representationslearning-based methods

Existing 3D representations—such as NeRFs, SDFs, voxels, point clouds, and octrees—are fragmented across generation and reconstruction tasks, lacking a unified, end-to-end evaluation framework. Method: We propose the first standardized benchmark for joint generation and reconstruction, covering the full pipeline: preprocessing, compression, generation, and reconstruction. Our framework systematically quantifies performance across reconstruction fidelity, generation quality, computational efficiency, and cross-scene generalization. Leveraging an autoencoder-based architecture, we enable fair, representation-agnostic comparison and uncover the dominant influence of reconstruction error on generative performance. Contribution/Results: We establish a practice-oriented 3D representation selection guide and release open-source code and benchmarks to foster community advancement. This work bridges critical gaps in holistic 3D representation evaluation and provides foundational insights into the interplay between reconstruction accuracy and generative capability.

Assessing joint impact of reconstruction errors on generation performanceComparing quality, efficiency, and generalization of 3D modelsEvaluating diverse 3D representations for reconstruction and generation

LATTICE: Democratize High-Fidelity 3D Generation at Scale

Nov 24, 2025
ZL
Zeqiang Lai
🏛️ MMLab, CUHK | Tencent Hunyuan

3D generative models face fundamental trade-offs between geometric fidelity, computational efficiency, and scalability due to the difficulty of modeling complex spatial structures and the high cost of volumetric representation. To address this, we propose VoxSet—a semi-structured voxel representation that explicitly encodes spatial topology under high compression, enabling position-aware generation and token-level test-time scaling. Our two-stage framework first generates sparse voxel anchors, then refines geometry to arbitrary resolution via a correction-flow Transformer. This unifies structured 3D modeling with efficient decoding, drastically reducing training overhead. Evaluated across multiple benchmarks, VoxSet achieves state-of-the-art performance in fidelity, diversity, and scalability metrics. It enables high-fidelity, large-scale, and flexible inference for 3D asset generation—supporting resolution-agnostic output and interactive editing without retraining.

Addressing computational complexity in 3D asset encodingBridging quality and scalability gaps in 3D generationEnabling efficient and structured high-fidelity 3D creation

Representation Learning for Point Cloud Understanding

Dec 05, 2025
SY
Siming Yan
🏛️ The University of Texas at Austin

This work addresses the limited representational capacity of point cloud learning by proposing a cross-modal representation learning framework that incorporates 2D visual priors. Methodologically, it innovatively leverages pre-trained 2D vision models to guide 3D point cloud network training via feature alignment—enabling effective 2D→3D knowledge transfer without naïve modality conversion. The framework integrates a point cloud encoder-decoder architecture, self-supervised pre-training, and primitive-level supervised segmentation into a unified learning paradigm. Extensive experiments on ScanNet and S3DIS benchmarks demonstrate substantial improvements in point cloud segmentation and scene understanding performance, validating the efficacy of 2D semantic priors in enhancing 3D representation learning. The proposed approach establishes a scalable, multimodal representation learning pathway for efficient 3D perception in autonomous driving and robotics applications.

Self-supervised learning methods for point cloud understandingSupervised learning for point cloud primitive segmentationTransfer learning from 2D to 3D using pre-trained models

Structured 3D Latents for Scalable and Versatile 3D Generation

Dec 02, 2024
JX
Jianfeng Xiang
🏛️ Tsinghua University | Microsoft Research | USTC

To address the demand for high-fidelity, diverse 3D asset generation and flexible editing, this paper introduces SLAT—a structured 3D implicit representation that jointly encodes sparse 3D mesh topology and multi-view visual foundation model features, enabling unified decoding into multiple 3D formats (e.g., radiance fields, 3D Gaussians, explicit meshes). Methodologically, SLAT pioneers a 2B-parameter Transformer architecture based on Rectified Flow for large-scale 3D latent-space modeling—the first of its kind. We curate a high-quality dataset of 500K 3D assets and perform end-to-end training. SLAT supports text- and image-conditioned generation, achieving state-of-the-art performance in fidelity, diversity, and editability. It enables real-time local 3D editing and dynamic output format switching. All code, models, and data are publicly released.

Develops a scalable 3D generation method for diverse assetsEnables high-quality text/image-conditioned 3D generation and local editingUnifies representation for multiple output formats like meshes and radiance fields

Latest Papers

What's happening recently
View more

The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.

3D visionbenchmark fragmentationdata representation

This work addresses the challenge of extracting semantically aligned cross fields from a single 2D image and robustly lifting them onto 3D mesh surfaces while preserving consistency in occluded regions. To this end, the authors propose CrossLift, a novel method that leverages text-to-image models to generate feature-aligned 2D quad images, from which pixel-wise directional fields are extracted. These fields are then lifted to the 3D surface through a two-stage interpolation scheme—combining intra-patch and cross-view weighted blending. A confidence-guided weighting mechanism is introduced to resolve directional ambiguities, enabling smooth extrapolation into occluded areas and supporting user sketch-based interaction. Experiments demonstrate that CrossLift produces quadrilateral meshes with superior semantic consistency on both organic and mechanical models, significantly outperforming existing approaches, and shows promise for applications in texture alignment and interactive design.

3D reconstructioncross fieldsquad meshing

This work addresses the challenges of high-resolution 3D medical image generation, where fully 3D models suffer from prohibitive computational costs and 2D slice-based approaches often produce anatomically inconsistent results. The authors propose an efficient and anatomically coherent 3D generative framework that decomposes volume synthesis into slice-wise generation coupled with inter-slice feature trajectory modeling. They introduce, for the first time in unconditional generation, a triplane drift loss to align the depth-trajectory distributions of real and generated volumes, and design a bidirectional z-context mixer to enhance inter-slice coherence. Built upon a 2D generator architecture, the method achieves superior single-slice image quality on BraTS 2023 and SynthRAD2023, near state-of-the-art performance in missing-modality reconstruction, approximately 135× lower inference cost, and significantly improved inter-slice consistency in MR-to-CT translation tasks.

2D-to-3D generation3D medical image generationanatomical consistency

Existing CAD learning approaches discretize B-Rep models into triangle meshes, thereby discarding the analytical surface representations and topological information essential for consistent instance-level analysis. This work proposes STEP-Parts, a deterministic pipeline that directly extracts geometric instance partitions from native STEP B-Rep data. The method defines partitions based on intrinsic B-Rep topology, merges faces using analytical surface types and near-tangent plane continuity criteria, and transfers labels to triangulated meshes via face-to-mesh correspondence mapping. STEP-Parts ensures boundary consistency across varying triangulations and processes the DeepCAD subset of the ABC dataset—comprising approximately 180,000 models—in under six hours. The resulting labels significantly enhance performance in implicit reconstruction-segmentation tasks and point cloud networks. Code and precomputed labels are publicly released.

Boundary RepresentationsCAD processinggeometric partitioning

This work addresses the inherent uncertainty in 3D scene reconstruction from limited observations—such as a single view, sparse pixels, or noisy images—by proposing a probabilistic framework that integrates Neural Radiance Fields (NeRF) with score-based diffusion models. The method represents the 3D scene as a stochastic latent variable, employs NeRF to model the likelihood of observations, and leverages a diffusion model to learn the prior distribution over the latent space. Crucially, it introduces, for the first time, a score-based diffusion mechanism to sample from the posterior distribution of the latent variables, enabling a unified treatment of uncertainty across diverse observation conditions. A two-stage training strategy jointly optimizes the reconstruction and prior networks, achieving high-fidelity 3D reconstructions under various settings—including single-view, multi-view, noisy images, sparse pixels, and depth inputs—while faithfully capturing task-specific uncertainties.

3D reconstructionobservation diversityposterior inference

Hot Scholars

JS

Jianbing Shen

Professor, University of Macau
Computer VisionMedical Image AnalysisVision and LanguageSelf-Driving Cars
KY

Kailun Yang

Professor. School of Artificial Intelligence and Robotics, Hunan University (HNU); KIT; UAH; ZJU
Computer VisionComputational OpticsIntelligent VehiclesAutonomous Driving
MJ

Minkyeong Jeon

Korea University
Representation LearningMulti-modal
SH

Sunghwan Hong

Postdoc @ ETHZ
Computer Vision3D VisionSfMSLAM
SK

Seungryong Kim

Associate Professor, KAIST
Computer VisionMachine Learning