promptable 3d instance segmentation

Designs, builds, or evaluates models and systems that take 3D data (e.g., point clouds, meshes, or voxels) and produce per‑instance segmentation masks for individual objects in a scene. These systems accept user prompts (points, clicks, boxes) or other interactive inputs to identify a target instance and support iterative, few‑click refinement and scalable inference across large 3D scans.

promptable3dinstancesegmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of accurately segmenting geometrically complex and materially diverse industrial objects, where linguistic or appearance-based prompts often fail. To overcome this limitation, the authors propose a geometry-conditioned prompting mechanism that leverages CAD models by introducing multi-view rendered images as geometric priors into the promptable segmentation framework SAM3. This approach enables appearance-agnostic, single-stage instance segmentation. By combining synthetic data training with multi-view CAD renderings, the method significantly improves segmentation accuracy and robustness across a wide range of materials and lighting conditions, effectively supporting industrial object segmentation tasks that cannot be reliably described by language or visual appearance alone.

CAD modelsgeometry-conditionedindustrial objects

OpenMaskDINO3D : Reasoning 3D Segmentation via Large Language Model

Jun 05, 2025
KZ
Kunshen Zhang
🏛️ Wuhan University

Current 3D instance segmentation methods rely on predefined categories or explicit manual prompts, limiting their ability to support open-vocabulary and natural-language-instruction-driven end-to-end reasoning. To address this, we propose the first open-vocabulary 3D instance segmentation framework for point clouds that operates without category priors or annotated prompts, directly interpreting implicit semantic intents from free-form language instructions. Our core innovations include a learnable SEG token and an object-identification mechanism, integrated within a unified architecture combining multimodal large language models, point cloud Transformers, and cross-modal alignment techniques to bridge the semantic gap between language and 3D geometry. Extensive experiments on ScanNet demonstrate that our method comprehensively outperforms prior approaches, achieving state-of-the-art performance across zero-shot transfer, few-shot segmentation, and complex instruction understanding tasks.

Lack of 3D reasoning segmentation frameworkNeed for implicit query text understanding in 3DPrecision challenges in 3D segmentation mask generation

Object Learning and Robust 3D Reconstruction

Apr 22, 2025
SS
Sara Sabour
🏛️ University of Toronto

This work addresses motion-induced artifacts in unsupervised 3D reconstruction of dynamic scenes caused by moving foreground objects. To tackle this, we propose a novel method that jointly models unsupervised 2D object decomposition and 3D geometric consistency. Our approach introduces FlowCapsules—a flow-guided network for unsupervised foreground segmentation—and a transient-object-mask-driven robust optimization kernel that detects and excludes dynamic objects under multi-view geometric consistency constraints. This kernel subsequently guides weighted bundle adjustment and NeRF training. Crucially, our method is the first to jointly learn object-centric representations and scene-level geometry without any annotations or controlled capture conditions. Experiments demonstrate significant improvements in both SfM and NeRF reconstruction accuracy on casually captured dynamic scenes, effectively eliminating motion artifacts. The framework advances explicit object-aware 3D understanding in open-world vision applications.

3D dynamic object detection via geometric consistencyRobust 3D modeling with transient object masksUnsupervised 2D object segmentation using motion cues

GeoMask3D: Geometrically Informed Mask Selection for Self-Supervised Point Cloud Learning in 3D

May 20, 2024
AB
Ali Bahri
🏛️ École de Technologie Supérieure

This work addresses the limitation of conventional random masking in point cloud self-supervised learning, which disregards intrinsic geometric structure and thus hinders representation learning. To this end, we propose GeoMask3D—a geometry-aware masking strategy. Its core innovations are twofold: (1) the first learnable geometric complexity metric, which guides the selection of mask regions based on structural intricacy; and (2) a feature-level full-partial knowledge distillation mechanism within a teacher-student framework, enabling context-guided geometric complexity prediction. GeoMask3D is seamlessly integrated into the masked autoencoder (MAE) paradigm, substantially enhancing the model’s capacity to capture fine-grained local geometry. Extensive experiments on benchmarks including ModelNet40 demonstrate state-of-the-art performance in both classification and few-shot recognition tasks, validating that geometry-aware masking delivers substantial and consistent gains for downstream performance.

Enhances self-supervised learning for 3D point cloudsFocuses on geometrically complex areas for robust featuresImproves performance in classification and few-shot tasks

To address the limitations of existing mesh part segmentation methods—such as reliance on category-specific annotations and poor generalization to unseen categories—this paper introduces the first zero-shot 3D mesh part segmentation framework. Methodologically, it jointly renders multi-view surface normals and Shape Diameter Functions (SDF) to generate 2D images, leverages the Segment Anything Model (SAM) to obtain cross-view 2D masks, and achieves part-level 3D segmentation via geometric-consistency-based mask aggregation and 2D-to-3D lifting. Key contributions include: (i) the first adaptation of a 2D vision foundation model for zero-shot transfer to 3D mesh segmentation; (ii) zero-training generalization to novel categories; and (iii) plug-and-play compatibility with upgraded SAM variants. Quantitative evaluation on standard and custom benchmarks shows segmentation accuracy competitive with or superior to conventional SDF-based methods. Human evaluation further confirms significantly improved semantic consistency across parts and stronger cross-shape generalization.

Generalization demonstrated via curated dataset and human evaluationMultimodal rendering and 2D-to-3D lifting for improved segmentationZero-shot mesh part segmentation overcoming existing limitations

Latest Papers

What's happening recently
View more

This work addresses the challenges of limited surface coverage, uneven point density, and severe class imbalance between planar and non-planar components in single LiDAR scans of architectural scenes. To this end, the authors propose an incidence-angle-aware geometric normalization sampling strategy that treats sampling as an active component of 3D segmentation. Under a fixed point budget, points are mapped into a normalized manifold space for voxel selection while preserving their original Euclidean coordinates for downstream learning. The method requires only point coordinates and normals and does not necessitate modifications to the network backbone. Evaluated on the SIP benchmark, it significantly improves average segmentation performance—particularly for non-planar structures such as ladders—and reduces sensitivity to sampling resolution.

3D segmentationclass imbalanceconstruction sites

This work addresses the challenge of generating semantic regions for 3D asset segmentation, which traditionally relies on manual intervention and struggles to integrate into interactive content creation pipelines. The authors propose a human-in-the-loop approach for producing editable semantic texture atlases by leveraging multi-view rendering and interactive 2D segmentation—combining SAM² with Label Studio—and back-projecting the results into UV space. A greedy set cover strategy is employed to select key views, enhancing computational efficiency. This method delivers the first unified, editable semantic atlas tailored for XR and game development workflows, enabling downstream tasks such as material assignment and style transfer. Experiments on eight cultural heritage objects demonstrate its effectiveness in handling complex geometries and accurately identifying fine details, cavities, and weak boundaries that require human refinement.

3D asset segmentationatlas generationhuman-in-the-loop

This work addresses the underexplored safety risks of image-to-3D generative models, which can be exploited to synthesize geometric structures posing real-world harm, while existing safeguards remain inadequate. The study introduces the first systematic definition and categorization of three types of harmful geometric content and proposes a multidimensional evaluation framework integrating geometric validity analysis, semantic scoring via multi-view vision-language models, human validation, and physical 3D printing tests. To assess robustness, the authors incorporate input perturbations and semantic obfuscation techniques. Experiments reveal that current models efficiently reconstruct harmful geometries, with commercial moderation systems detecting fewer than 0.3% of such instances. The proposed stacked defense mechanism reduces harmful content retention to below 1%, albeit with an 11% false positive rate, thereby substantially enhancing overall safety.

3D content safetyadversarial misuseharmful geometry

This work addresses the challenge of efficiently reconstructing structured and editable CAD models from unorganized point clouds by proposing an end-to-end deep learning framework. The method innovatively introduces an extrusion-based segmentation strategy that decomposes input point clouds into individual extrusion primitives. This approach significantly enhances data diversity and reconstruction robustness without increasing model complexity. Experimental results demonstrate that the proposed framework substantially improves both the accuracy and generalization capability of CAD model reconstruction, offering a more reliable automated solution for applications in reverse engineering and manufacturing quality control.

CAD reconstructionextrusion segmentationpoint cloud

This work addresses the integration bias in existing multi-pretrained-model approaches to 3D instance segmentation, which arises from discrepancies in model confidence scores and degrades segmentation accuracy. To overcome this limitation, the authors propose a training-free 3D instance segmentation method that introduces, for the first time, a geometry–vision correspondence mechanism. This mechanism achieves precise alignment between 3D geometric cues and 2D visual cues, effectively eliminating confidence bias. By integrating 3D proposal generation with mask-aware CLIP feature extraction, the method enables high-quality, training-free model ensembling. The approach achieves state-of-the-art performance on multiple established 3D segmentation benchmarks and demonstrates exceptional generalization capability in open-vocabulary semantic segmentation tasks.

3D instance segmentationconfidence biasensemble learning

Hot Scholars

MR

Martin R. Oswald

University of Amsterdam
3D Computer VisionRepresentation LearningApplied Machine LearningOptimization
SP

Szymon Płotka

Jagiellonian University
Machine LearningDeep LearningComputer VisionMedical Imaging
AS

Arkadiusz Sitek

Massachusetts General Hospital, Harvard Medical School
HealthcareMachine LearningMedical Physics
TG

Theo Gevers

Computer Vision Research Group, University of Amsterdam (UvA); 3DUniversum
computer visionimage understandingdeep learningobject recognition
CL

Chenxin Li

The Chinese University of Hong Kong
Multimodal LLMAgentWorld Model