multi-view mask reprojection

Designs and implements methods that reproject 2D segmentation masks from multiple calibrated views onto 3D geometry (meshes or surfaces) using geometry-aware ray casting, associating mask labels to mesh faces, edges, or surface points. Builds fusion and lifting pipelines that aggregate per-view masks into consistent per-surface or per-part segmentations, producing fused mask-attributed meshes or per-part surface point clouds.

multi-viewmaskreprojection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the limitations of existing mesh part segmentation methods—such as reliance on category-specific annotations and poor generalization to unseen categories—this paper introduces the first zero-shot 3D mesh part segmentation framework. Methodologically, it jointly renders multi-view surface normals and Shape Diameter Functions (SDF) to generate 2D images, leverages the Segment Anything Model (SAM) to obtain cross-view 2D masks, and achieves part-level 3D segmentation via geometric-consistency-based mask aggregation and 2D-to-3D lifting. Key contributions include: (i) the first adaptation of a 2D vision foundation model for zero-shot transfer to 3D mesh segmentation; (ii) zero-training generalization to novel categories; and (iii) plug-and-play compatibility with upgraded SAM variants. Quantitative evaluation on standard and custom benchmarks shows segmentation accuracy competitive with or superior to conventional SDF-based methods. Human evaluation further confirms significantly improved semantic consistency across parts and stronger cross-shape generalization.

Generalization demonstrated via curated dataset and human evaluationMultimodal rendering and 2D-to-3D lifting for improved segmentationZero-shot mesh part segmentation overcoming existing limitations

View2CAD: Reconstructing View-Centric CAD Models from Single RGB-D Scans

Apr 05, 2025
JN
James Noeckel
🏛️ University of Washington | Massachusetts Institute of Technology

This work addresses the challenging problem of reconstructing boundary-representation (B-rep) CAD models from a single RGB-D image under viewpoint-centered observation. We propose View-based B-Rep (VB-Rep), the first view-aware B-rep representation that explicitly encodes visibility constraints and geometric uncertainty—overcoming the limitation of existing methods that require complete, noise-free 3D inputs. Our method integrates panoramic image segmentation with depth-aware iterative geometric optimization, leveraging semantic priors to guide accurate boundary reconstruction and suppress geometric hallucinations and topological inaccuracies under partial observability. The output is a fully parametric, editable, and manufacturable B-rep model. Evaluated on both synthetic and real-world RGB-D datasets, our approach achieves significantly higher reconstruction fidelity, effectively bridging the semantic and geometric gap between real-world scenes and parametric CAD modeling.

Address partial geometry observation challengesIntroduce view-centric B-rep for uncertainty handlingReconstruct CAD models from single RGB-D scans

3D Mesh Editing using Masked LRMs

Dec 11, 2024
WG
Will Gao
🏛️ University of Chicago | Meta Reality Labs

This work addresses the problem of localized 3D mesh editing guided by a single edited image. We propose an interactive editing framework based on mask-conditioned reconstruction: user-specified 3D regions serve as geometric masks, and the edited image acts as a conditioning signal to guide a Large Reconstruction Model (LRM) to reconstruct only the masked regions while preserving high fidelity in the unmasked areas. To our knowledge, this is the first method to adapt LRM for real-time, mask-conditioned mesh editing. Our approach integrates multi-view-consistent mask rendering, stochastic 3D occlusion synthesis, and single-view conditional injection, enabling high-quality geometric updates in a single forward pass. The framework supports diverse semantic edits—including deformation, part replacement, and detail sculpting—achieving state-of-the-art reconstruction quality while accelerating inference by 10× over the best prior baseline.

Editing 3D shapes via conditional reconstructionGenerating geometry in masked regions from imagesPerforming mesh edits faster than prior methods

Existing zero-shot 3D instance segmentation methods often yield fragmented results due to their neglect of multi-view consistency and 3D geometric priors. This work proposes a coarse-to-fine zero-shot segmentation framework that first leverages coarse 3D fragments as a shared reference for cross-view matching, enhancing consistency through 3D-guided multi-view 2D mask alignment. To address occlusion ambiguities, a depth-consistency weighting mechanism is introduced, which, combined with SAM-generated mask fusion and reliability assessment of 3D-to-2D projections, improves instance completeness. The proposed method significantly outperforms current approaches across multiple benchmarks—including ScanNetV2, ScanNet200, ScanNet++, Replica, and Matterport3D—delivering more complete and robust zero-shot 3D instance segmentation.

3D priors3D-to-2D correspondencemask consistency

Pixels2Points: Fusing 2D and 3D Features for Facial Skin Segmentation

Apr 28, 2025
VY
Victoria Yue Chen
🏛️ ETH Zürich | Google

To address inaccurate skin-region segmentation in 3D facial scans—which degrades registration accuracy—this paper proposes an end-to-end, mesh-level segmentation method that jointly leverages multi-view 2D semantic features and 3D geometric features. Innovatively, it freezes a Vision Transformer (ViT) backbone to extract robust 2D semantic features, then aligns and fuses them into mesh vertices via learnable 3D feature projection and voxel-based aggregation. Subsequent refinement is performed using a graph convolutional network (GCN). Notably, the method requires no ground-truth skin annotations and achieves strong cross-domain generalization when trained exclusively on synthetic data. Evaluated on real-world 3D facial scans, it improves registration accuracy by 8.89% over pure 2D baselines and by 14.3% over pure 3D baselines, significantly enhancing facial registration quality.

Accurate skin segmentation on 3D head scansFusion of 2D image and 3D geometric featuresImproving face registration quality by 8.89-14.3%

Latest Papers

What's happening recently
View more

This work addresses the challenge of degraded 3D surface reconstruction accuracy caused by missing geometric information in LiDAR point clouds due to limited scanning range and occlusions. To tackle this issue, the authors propose a reconstruction method based on plane classification and priority-driven growth. The approach categorizes scene planes into three visibility classes—highly visible, partially visible, and invisible—and employs a hierarchical spatial partitioning scheme. Coupled with a min-cut optimization strategy, it generates compact, watertight polygonal models that effectively recover missing geometric details. Evaluated on public datasets, the method significantly outperforms current state-of-the-art techniques, achieving higher reconstruction fidelity while preserving model compactness.

3D visionLiDAR scanningmissing details

Hot Scholars

MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
YQ

Yue Qi

Beihang University
JK

Jan Kautz

Vice President of Research, NVIDIA Research
Computer VisionMachine LearningVisual Computing