Score
Designs, builds, or evaluates models and systems that take 3D data (e.g., point clouds, meshes, or voxels) and produce per‑instance segmentation masks for individual objects in a scene. These systems accept user prompts (points, clicks, boxes) or other interactive inputs to identify a target instance and support iterative, few‑click refinement and scalable inference across large 3D scans.
Point cloud data pose unique challenges for deep learning due to their unordered and irregular structure, as well as sensor noise and occlusion. This work systematically surveys representative models for point cloud classification, part segmentation, and semantic segmentation from the perspective of backbone architectures, synthesizing key technical approaches including ordering-based transformation, local geometric feature extraction, permutation-invariant processing, and self-attention mechanisms. The study provides an in-depth analysis of their architectural characteristics, innovative designs, and inherent limitations. Through comprehensive performance comparisons on mainstream benchmarks, it clarifies the evolutionary trajectory of deep learning methods in 3D vision tasks, offering a clear reference for model selection and future research directions.
This work addresses the challenge of accurately segmenting geometrically complex and materially diverse industrial objects, where linguistic or appearance-based prompts often fail. To overcome this limitation, the authors propose a geometry-conditioned prompting mechanism that leverages CAD models by introducing multi-view rendered images as geometric priors into the promptable segmentation framework SAM3. This approach enables appearance-agnostic, single-stage instance segmentation. By combining synthetic data training with multi-view CAD renderings, the method significantly improves segmentation accuracy and robustness across a wide range of materials and lighting conditions, effectively supporting industrial object segmentation tasks that cannot be reliably described by language or visual appearance alone.
Current 3D instance segmentation methods rely on predefined categories or explicit manual prompts, limiting their ability to support open-vocabulary and natural-language-instruction-driven end-to-end reasoning. To address this, we propose the first open-vocabulary 3D instance segmentation framework for point clouds that operates without category priors or annotated prompts, directly interpreting implicit semantic intents from free-form language instructions. Our core innovations include a learnable SEG token and an object-identification mechanism, integrated within a unified architecture combining multimodal large language models, point cloud Transformers, and cross-modal alignment techniques to bridge the semantic gap between language and 3D geometry. Extensive experiments on ScanNet demonstrate that our method comprehensively outperforms prior approaches, achieving state-of-the-art performance across zero-shot transfer, few-shot segmentation, and complex instruction understanding tasks.
This work addresses motion-induced artifacts in unsupervised 3D reconstruction of dynamic scenes caused by moving foreground objects. To tackle this, we propose a novel method that jointly models unsupervised 2D object decomposition and 3D geometric consistency. Our approach introduces FlowCapsules—a flow-guided network for unsupervised foreground segmentation—and a transient-object-mask-driven robust optimization kernel that detects and excludes dynamic objects under multi-view geometric consistency constraints. This kernel subsequently guides weighted bundle adjustment and NeRF training. Crucially, our method is the first to jointly learn object-centric representations and scene-level geometry without any annotations or controlled capture conditions. Experiments demonstrate significant improvements in both SfM and NeRF reconstruction accuracy on casually captured dynamic scenes, effectively eliminating motion artifacts. The framework advances explicit object-aware 3D understanding in open-world vision applications.
This work addresses the limitation of conventional random masking in point cloud self-supervised learning, which disregards intrinsic geometric structure and thus hinders representation learning. To this end, we propose GeoMask3D—a geometry-aware masking strategy. Its core innovations are twofold: (1) the first learnable geometric complexity metric, which guides the selection of mask regions based on structural intricacy; and (2) a feature-level full-partial knowledge distillation mechanism within a teacher-student framework, enabling context-guided geometric complexity prediction. GeoMask3D is seamlessly integrated into the masked autoencoder (MAE) paradigm, substantially enhancing the model’s capacity to capture fine-grained local geometry. Extensive experiments on benchmarks including ModelNet40 demonstrate state-of-the-art performance in both classification and few-shot recognition tasks, validating that geometry-aware masking delivers substantial and consistent gains for downstream performance.
To address the limitations of existing mesh part segmentation methods—such as reliance on category-specific annotations and poor generalization to unseen categories—this paper introduces the first zero-shot 3D mesh part segmentation framework. Methodologically, it jointly renders multi-view surface normals and Shape Diameter Functions (SDF) to generate 2D images, leverages the Segment Anything Model (SAM) to obtain cross-view 2D masks, and achieves part-level 3D segmentation via geometric-consistency-based mask aggregation and 2D-to-3D lifting. Key contributions include: (i) the first adaptation of a 2D vision foundation model for zero-shot transfer to 3D mesh segmentation; (ii) zero-training generalization to novel categories; and (iii) plug-and-play compatibility with upgraded SAM variants. Quantitative evaluation on standard and custom benchmarks shows segmentation accuracy competitive with or superior to conventional SDF-based methods. Human evaluation further confirms significantly improved semantic consistency across parts and stronger cross-shape generalization.
This work addresses the challenges of limited surface coverage, uneven point density, and severe class imbalance between planar and non-planar components in single LiDAR scans of architectural scenes. To this end, the authors propose an incidence-angle-aware geometric normalization sampling strategy that treats sampling as an active component of 3D segmentation. Under a fixed point budget, points are mapped into a normalized manifold space for voxel selection while preserving their original Euclidean coordinates for downstream learning. The method requires only point coordinates and normals and does not necessitate modifications to the network backbone. Evaluated on the SIP benchmark, it significantly improves average segmentation performance—particularly for non-planar structures such as ladders—and reduces sensitivity to sampling resolution.
This work addresses the challenge of generating semantic regions for 3D asset segmentation, which traditionally relies on manual intervention and struggles to integrate into interactive content creation pipelines. The authors propose a human-in-the-loop approach for producing editable semantic texture atlases by leveraging multi-view rendering and interactive 2D segmentation—combining SAM² with Label Studio—and back-projecting the results into UV space. A greedy set cover strategy is employed to select key views, enhancing computational efficiency. This method delivers the first unified, editable semantic atlas tailored for XR and game development workflows, enabling downstream tasks such as material assignment and style transfer. Experiments on eight cultural heritage objects demonstrate its effectiveness in handling complex geometries and accurately identifying fine details, cavities, and weak boundaries that require human refinement.
This work addresses the underexplored safety risks of image-to-3D generative models, which can be exploited to synthesize geometric structures posing real-world harm, while existing safeguards remain inadequate. The study introduces the first systematic definition and categorization of three types of harmful geometric content and proposes a multidimensional evaluation framework integrating geometric validity analysis, semantic scoring via multi-view vision-language models, human validation, and physical 3D printing tests. To assess robustness, the authors incorporate input perturbations and semantic obfuscation techniques. Experiments reveal that current models efficiently reconstruct harmful geometries, with commercial moderation systems detecting fewer than 0.3% of such instances. The proposed stacked defense mechanism reduces harmful content retention to below 1%, albeit with an 11% false positive rate, thereby substantially enhancing overall safety.
This work addresses the challenge of efficiently reconstructing structured and editable CAD models from unorganized point clouds by proposing an end-to-end deep learning framework. The method innovatively introduces an extrusion-based segmentation strategy that decomposes input point clouds into individual extrusion primitives. This approach significantly enhances data diversity and reconstruction robustness without increasing model complexity. Experimental results demonstrate that the proposed framework substantially improves both the accuracy and generalization capability of CAD model reconstruction, offering a more reliable automated solution for applications in reverse engineering and manufacturing quality control.
This work addresses the integration bias in existing multi-pretrained-model approaches to 3D instance segmentation, which arises from discrepancies in model confidence scores and degrades segmentation accuracy. To overcome this limitation, the authors propose a training-free 3D instance segmentation method that introduces, for the first time, a geometry–vision correspondence mechanism. This mechanism achieves precise alignment between 3D geometric cues and 2D visual cues, effectively eliminating confidence bias. By integrating 3D proposal generation with mask-aware CLIP feature extraction, the method enables high-quality, training-free model ensembling. The approach achieves state-of-the-art performance on multiple established 3D segmentation benchmarks and demonstrates exceptional generalization capability in open-vocabulary semantic segmentation tasks.