Score
Designs and implements methods that reproject 2D segmentation masks from multiple calibrated views onto 3D geometry (meshes or surfaces) using geometry-aware ray casting, associating mask labels to mesh faces, edges, or surface points. Builds fusion and lifting pipelines that aggregate per-view masks into consistent per-surface or per-part segmentations, producing fused mask-attributed meshes or per-part surface point clouds.
To address the limitations of existing mesh part segmentation methods—such as reliance on category-specific annotations and poor generalization to unseen categories—this paper introduces the first zero-shot 3D mesh part segmentation framework. Methodologically, it jointly renders multi-view surface normals and Shape Diameter Functions (SDF) to generate 2D images, leverages the Segment Anything Model (SAM) to obtain cross-view 2D masks, and achieves part-level 3D segmentation via geometric-consistency-based mask aggregation and 2D-to-3D lifting. Key contributions include: (i) the first adaptation of a 2D vision foundation model for zero-shot transfer to 3D mesh segmentation; (ii) zero-training generalization to novel categories; and (iii) plug-and-play compatibility with upgraded SAM variants. Quantitative evaluation on standard and custom benchmarks shows segmentation accuracy competitive with or superior to conventional SDF-based methods. Human evaluation further confirms significantly improved semantic consistency across parts and stronger cross-shape generalization.
This work addresses the challenging problem of reconstructing boundary-representation (B-rep) CAD models from a single RGB-D image under viewpoint-centered observation. We propose View-based B-Rep (VB-Rep), the first view-aware B-rep representation that explicitly encodes visibility constraints and geometric uncertainty—overcoming the limitation of existing methods that require complete, noise-free 3D inputs. Our method integrates panoramic image segmentation with depth-aware iterative geometric optimization, leveraging semantic priors to guide accurate boundary reconstruction and suppress geometric hallucinations and topological inaccuracies under partial observability. The output is a fully parametric, editable, and manufacturable B-rep model. Evaluated on both synthetic and real-world RGB-D datasets, our approach achieves significantly higher reconstruction fidelity, effectively bridging the semantic and geometric gap between real-world scenes and parametric CAD modeling.
This work addresses the problem of localized 3D mesh editing guided by a single edited image. We propose an interactive editing framework based on mask-conditioned reconstruction: user-specified 3D regions serve as geometric masks, and the edited image acts as a conditioning signal to guide a Large Reconstruction Model (LRM) to reconstruct only the masked regions while preserving high fidelity in the unmasked areas. To our knowledge, this is the first method to adapt LRM for real-time, mask-conditioned mesh editing. Our approach integrates multi-view-consistent mask rendering, stochastic 3D occlusion synthesis, and single-view conditional injection, enabling high-quality geometric updates in a single forward pass. The framework supports diverse semantic edits—including deformation, part replacement, and detail sculpting—achieving state-of-the-art reconstruction quality while accelerating inference by 10× over the best prior baseline.
Existing zero-shot 3D instance segmentation methods often yield fragmented results due to their neglect of multi-view consistency and 3D geometric priors. This work proposes a coarse-to-fine zero-shot segmentation framework that first leverages coarse 3D fragments as a shared reference for cross-view matching, enhancing consistency through 3D-guided multi-view 2D mask alignment. To address occlusion ambiguities, a depth-consistency weighting mechanism is introduced, which, combined with SAM-generated mask fusion and reliability assessment of 3D-to-2D projections, improves instance completeness. The proposed method significantly outperforms current approaches across multiple benchmarks—including ScanNetV2, ScanNet200, ScanNet++, Replica, and Matterport3D—delivering more complete and robust zero-shot 3D instance segmentation.
To address inaccurate skin-region segmentation in 3D facial scans—which degrades registration accuracy—this paper proposes an end-to-end, mesh-level segmentation method that jointly leverages multi-view 2D semantic features and 3D geometric features. Innovatively, it freezes a Vision Transformer (ViT) backbone to extract robust 2D semantic features, then aligns and fuses them into mesh vertices via learnable 3D feature projection and voxel-based aggregation. Subsequent refinement is performed using a graph convolutional network (GCN). Notably, the method requires no ground-truth skin annotations and achieves strong cross-domain generalization when trained exclusively on synthetic data. Evaluated on real-world 3D facial scans, it improves registration accuracy by 8.89% over pure 2D baselines and by 14.3% over pure 3D baselines, significantly enhancing facial registration quality.
本文提出AnyGS2Mesh,一种前馈框架,直接从3D高斯点阵表示中重建3D网格,解决现有方法依赖迭代优化导致的速度慢和分辨率限制问题。
本文提出VGGT-CAD,通过几何先验和多视图特征融合解决参数化CAD模型从有限视角重建的问题。
本文提出ReconSplat模型,通过结合3D高斯点云和多视图潜在扩散模型解决未观测区域的合理视角生成与几何一致性问题。
本文提出SAM3D-Part,通过用户提示从3D对象中选择并生成特定部分,解决现有方法无法按需生成及准确对齐的问题。
This work addresses the challenge of degraded 3D surface reconstruction accuracy caused by missing geometric information in LiDAR point clouds due to limited scanning range and occlusions. To tackle this issue, the authors propose a reconstruction method based on plane classification and priority-driven growth. The approach categorizes scene planes into three visibility classes—highly visible, partially visible, and invisible—and employs a hierarchical spatial partitioning scheme. Coupled with a min-cut optimization strategy, it generates compact, watertight polygonal models that effectively recover missing geometric details. Evaluated on public datasets, the method significantly outperforms current state-of-the-art techniques, achieving higher reconstruction fidelity while preserving model compactness.