3d reconstruction

Designs, builds, or evaluates algorithms and systems that recover three-dimensional geometry and appearance from images or sensor data, producing meshes, volumetric or surface representations, UV textures, and object pose (including 6-DoF placement). Work covers monocular/single-view, multi-view and multi-view-stereo methods, object-centric and rigid-object extraction, self-supervised view/frame consistency, and alignment of reconstructed geometry to semantic models.

3dreconstruction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Physically Compatible 3D Object Modeling from a Single Image

May 30, 2024
MG
Minghao Guo
🏛️ MIT | UMass Amherst

This work addresses the lack of physical plausibility in single-image 3D reconstruction. We propose the first physics-compatible reconstruction framework that enforces static equilibrium as a hard constraint. Methodologically, we explicitly decouple and jointly optimize material stiffness, external loading forces, and the static equilibrium geometry; deformation responses are modeled via differentiable physics simulation, enabling gradient-based joint optimization of all variables. Our approach breaks from conventional simplifications—such as rigid-body assumptions or neglect of external forces—by embedding real-world physical constraints directly into the single-image reconstruction pipeline. Evaluated on Objaverse, our method yields reconstructions with significantly improved mechanical stability, suitable for downstream dynamic simulation and 3D printing. Physical validation via real-world force testing further confirms the structural robustness of the generated models.

3D modelingmaterial propertiesphysical stability

This work proposes a fully automated method for reconstructing high-fidelity three-dimensional models from multiple orthogonal views of an object. The approach begins by extracting key control points using Harris corner detection, followed by generating mutually orthogonal bounding volumes through orthographic projection and constructing a 3D point cloud from their intersections. Subsequently, computational geometry algorithms are employed to recover the surface topology, and the resulting model is rendered using OpenGL for visualization. The entire pipeline operates without human intervention, achieving end-to-end reconstruction from 2D orthogonal projections to a complete 3D structure. The method demonstrates notable advantages in geometric consistency and reconstruction completeness compared to existing approaches.

3D modeling3D reconstructionautomated modeling

Existing multi-view human mesh reconstruction methods rely heavily on cumbersome camera calibration or multi-view training data, limiting their generalization capability. This work proposes a training-free test-time optimization framework that, for the first time, leverages pre-trained single-view human mesh recovery models (e.g., HMR) as strong priors, integrating multi-view consistency and anatomical constraints to achieve high-fidelity, calibration-free reconstruction under arbitrary camera configurations. By eliminating the need for multi-view supervised training, the method significantly enhances generalization and achieves performance on par with or superior to state-of-the-art approaches that explicitly require multi-view training data, as demonstrated on standard benchmarks.

camera calibrationgeneralizationmulti-view human mesh recovery

View2CAD: Reconstructing View-Centric CAD Models from Single RGB-D Scans

Apr 05, 2025
JN
James Noeckel
🏛️ University of Washington | Massachusetts Institute of Technology

This work addresses the challenging problem of reconstructing boundary-representation (B-rep) CAD models from a single RGB-D image under viewpoint-centered observation. We propose View-based B-Rep (VB-Rep), the first view-aware B-rep representation that explicitly encodes visibility constraints and geometric uncertainty—overcoming the limitation of existing methods that require complete, noise-free 3D inputs. Our method integrates panoramic image segmentation with depth-aware iterative geometric optimization, leveraging semantic priors to guide accurate boundary reconstruction and suppress geometric hallucinations and topological inaccuracies under partial observability. The output is a fully parametric, editable, and manufacturable B-rep model. Evaluated on both synthetic and real-world RGB-D datasets, our approach achieves significantly higher reconstruction fidelity, effectively bridging the semantic and geometric gap between real-world scenes and parametric CAD modeling.

Address partial geometry observation challengesIntroduce view-centric B-rep for uncertainty handlingReconstruct CAD models from single RGB-D scans

This work addresses the challenge of degraded 3D surface reconstruction accuracy caused by missing geometric information in LiDAR point clouds due to limited scanning range and occlusions. To tackle this issue, the authors propose a reconstruction method based on plane classification and priority-driven growth. The approach categorizes scene planes into three visibility classes—highly visible, partially visible, and invisible—and employs a hierarchical spatial partitioning scheme. Coupled with a min-cut optimization strategy, it generates compact, watertight polygonal models that effectively recover missing geometric details. Evaluated on public datasets, the method significantly outperforms current state-of-the-art techniques, achieving higher reconstruction fidelity while preserving model compactness.

3D visionLiDAR scanningmissing details

Latest Papers

What's happening recently
View more

This study addresses the lack of a unified 3D reconstruction framework in manufacturing, particularly under challenging conditions involving reflective surfaces and dynamic environments. Through a systematic review of 106 publications, the work proposes a structured classification framework tailored to manufacturing scenarios, organizing techniques into three stages: data acquisition, point cloud generation and post-processing, and application. It integrates non-contact methods—such as structured light and stereo vision—with deep learning–driven feature extraction strategies. The analysis reveals that 40% of applications focus on quality inspection, with existing approaches achieving sub-millimeter accuracy in controlled settings. The study identifies multi-sensor fusion and hybrid systems as critical future directions and highlights several research gaps and technical challenges that warrant further investigation.

3D reconstructiondynamic environmentsmanufacturing

This work addresses the limitations of existing 3D inpainting methods, which predominantly focus on geometric completion while neglecting texture recovery and struggling with complex objects. To overcome these challenges, we propose the first end-to-end framework that jointly reconstructs both shape and texture of damaged objects from multi-view images. Our approach leverages automatically synthesized damaged-complete data pairs, a Mask Self-Perceiver module, and a Depth-Aware Mask Rectifier, integrated within a coarse-to-fine 3D mesh reconstruction strategy. This enables high-resolution, semantically consistent, and view-coherent inpainting results. Extensive evaluations on both synthetic and real-world benchmarks demonstrate that our method significantly outperforms state-of-the-art techniques in multi-view image inpainting and textured 3D reconstruction.

3D object restorationbroken objectsmulti-view reconstruction

The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.

3D visionbenchmark fragmentationdata representation

Existing image-based reconstruction methods struggle to produce explicit, watertight, and smooth geometries suitable for numerical simulation. This work proposes a unified optimization framework that directly constructs multi-patch B-spline boundary representations from sparse RGB images, simultaneously achieving geometric reconstruction and simulation-ready modeling. By projecting observed fields—such as temperature or semantic labels—onto the same B-spline basis, the approach seamlessly integrates visual reconstruction with isogeometric analysis (IGA). For the first time, it end-to-end combines image-driven geometry generation with physical field mapping, enabling high-fidelity thermal simulation and modal analysis. The method offers a practical solution for digital twin applications and simulation-driven design.

boundary representationgeometric modelingimage-based reconstruction

This work addresses the challenge of sim-to-real transfer in industrial visual inspection, where multiple domain gaps arise from discrepancies in sensors, lighting conditions, materials, and defect patterns. The authors propose a unified framework centered on the availability of prior knowledge, systematically integrating three scenarios: CAD-available, CAD-unavailable, and boundary cases, thereby harmonizing CAD-driven pose estimation and CAD-free anomaly detection paradigms. Their approach combines CAD-based rendering, RGB-D simulation, synthetic defect generation, pre-trained features, vision-language priors, and test-time geometric consistency verification. Experiments on T-LESS/BOP, MVTec AD, and VisA benchmarks demonstrate that transfer performance hinges more critically on source distribution design, detector capacity, and minimal real-world calibration than on the quantity of CAD renderings; notably, CAD models at test time effectively enable mask generation, pose refinement, and depth consistency validation.

CAD availabilitydomain gapindustrial visual inspection

Hot Scholars

AV

Andrea Vedaldi

University of Oxford
Computer VisionMachine Learning
WZ

Wenzhao Zheng

EECS, University of California, Berkeley
Large ModelsEmbodied AgentsAutonomous Driving
MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
XZ

Xiaowei Zhou

Professor of Computer Science, Zhejiang University
Computer VisionComputer Graphics
SW

Shenlong Wang

University of Illinois at Urbana-Champaign
Computer VisionRobot PerceptionAutonomous Driving