Score
Designs and implements models and end-to-end pipelines that detect and delineate individual object instances in 2D imagery and 3D point-cloud/PLY data, producing pixel- or point-level masks, separating overlapping instances, extracting regions of interest, and visualizing instance saliency (including adapting region-based architectures like Mask R-CNN). Builds and runs segmentation benchmarking and evaluation workflows—selecting and computing segmentation metrics, training and validating instance-aware labeling and ply/point distinction methods, and ensuring robustness to noisy or blurred boundaries.
Convolutional Neural Networks (CNNs) exhibit limitations in modeling long-range dependencies, multi-scale objects, and contextual information for image segmentation. Method: This paper presents a systematic survey of segmentation-specific Transformer architectures. It introduces a unified architectural taxonomy covering encoder-decoder designs, multi-scale feature fusion strategies, mask prediction heads, and attention variants—including windowed and axial attention. A “challenge–solution” mapping framework is established to identify three key bottlenecks: computational overhead, poor generalization under low-data regimes, and constraints on real-time deployment. Contribution/Results: The survey proposes two principal evolutionary directions—lightweight design and data-efficient learning—and establishes a unified evaluation benchmark to delineate state-of-the-art performance boundaries. Collectively, this work provides a principled, industrially viable roadmap for deploying Transformer-based segmentation systems.
Unsupervised 2D instance segmentation struggles with disentangling overlapping objects, as existing approaches model only semantic information and lack spatial decoupling capability. To address this, we propose the first point-cloud-based 3D semantic cutting paradigm: leveraging scene-level point clouds to construct a geometry-aware 3D spatial representation, projecting 2D semantic masks into 3D space and decoupling them into precise instances. We design a spatial importance function to enhance boundary semantics in 3D and introduce a three-component spatial confidence mechanism to mitigate ambiguity in pseudo-labels. Our method integrates point-cloud geometric modeling, joint training with a class-agnostic detector, and pseudo-label confidence refinement. Evaluated on PASCAL VOC and COCO benchmarks, our approach significantly outperforms state-of-the-art unsupervised instance segmentation and object detection methods.
This work addresses the challenge of instance segmentation in low-texture, uniformly colored industrial scenes, where existing methods struggle due to insufficient texture and contextual cues, and where instance definitions often vary by task. The authors propose a boundary-aware few-shot instance segmentation framework that leverages a pretrained foundation model to extract visual features and employs a lightweight signed distance function (SDF) head to predict boundary-sensitive distance maps. Instance masks are then reconstructed from these SDF predictions. By shifting supervision from internal appearance to explicit boundary modeling, the approach enables flexible instance definition—such as whole objects or subparts—via mask conditioning. A pixel-wise shallow MLP facilitates rapid training and fine-grained control. Experiments demonstrate superior few-shot generalization, robustness, and precise control over segmentation granularity on low-texture industrial components and food imagery.
This work addresses motion-induced artifacts in unsupervised 3D reconstruction of dynamic scenes caused by moving foreground objects. To tackle this, we propose a novel method that jointly models unsupervised 2D object decomposition and 3D geometric consistency. Our approach introduces FlowCapsules—a flow-guided network for unsupervised foreground segmentation—and a transient-object-mask-driven robust optimization kernel that detects and excludes dynamic objects under multi-view geometric consistency constraints. This kernel subsequently guides weighted bundle adjustment and NeRF training. Crucially, our method is the first to jointly learn object-centric representations and scene-level geometry without any annotations or controlled capture conditions. Experiments demonstrate significant improvements in both SfM and NeRF reconstruction accuracy on casually captured dynamic scenes, effectively eliminating motion artifacts. The framework advances explicit object-aware 3D understanding in open-world vision applications.
This work proposes L2G-Det, a novel framework for detecting and segmenting novel object instances in open-world scenarios using only a few template images, addressing challenges such as occlusion and background clutter. Departing from conventional proposal-based paradigms, L2G-Det introduces an end-to-end local-to-global pipeline: it generates candidate points through dense local patch matching between template and query images, then leverages these refined candidates to guide an enhanced Segment Anything Model (SAM) in reconstructing complete instance masks. Key innovations include the introduction of instance-level object tokens to strengthen SAM’s instance awareness and a tailored instance-specific prompting mechanism. Experiments demonstrate that L2G-Det significantly outperforms existing proposal-based methods in complex open-world settings, achieving markedly improved robustness and accuracy in novel instance detection and segmentation.
This study systematically evaluates the robustness and generalization of instance segmentation models under realistic image degradations (15 types of noise, blur, etc.) and cross-domain data (diverse acquisition settings). Leveraging benchmarks including COCO, we establish a multi-dimensional evaluation framework to comparatively analyze mainstream architectures, normalization strategies (Group Normalization vs. Batch Normalization), pretraining paradigms, and single- versus multi-stage detection frameworks. Key findings include: (i) Group Normalization significantly enhances degradation robustness (+8.2% mAP), whereas Batch Normalization improves cross-domain generalization (+5.7% mAP); (ii) multi-stage detectors exhibit scale-invariant performance, while single-stage models suffer from poor resolution generalization. We release a reproducible robustness leaderboard and an evidence-based model selection guide, providing critical insights for deploying instance segmentation models in real-world scenarios and informing architectural design and pretraining strategy choices.
This work addresses the limited generalization of tree instance segmentation in forest point clouds across diverse sensors, platforms, and forest types by proposing a unified, sensor- and platform-agnostic framework. Built upon the Point Transformer v3 backbone, the method integrates a lightweight semantic head with a tree-focused cross-attention mask decoder, enhanced by tree-aware query initialization, one-to-many seed supervision, and an asymmetric mask scoring mechanism to significantly improve instance separation accuracy in dense stands. The study also introduces FOR-instance v3, a large-scale benchmark dataset encompassing diverse ecosystems. Evaluated on the FOR-instance v2 test set, the approach achieves 90.5% precision, 80.2% recall, 85.0% F1 score, 90.7% coverage, and 87.6% semantic mIoU, substantially outperforming existing methods and demonstrating exceptional cross-domain generalization capability.
This work addresses the challenges of category-agnostic 3D instance segmentation in unknown environments, where frame-by-frame projection often leads to identity fragmentation and spatial discontinuities. To overcome these limitations without requiring any 3D training data, the authors propose a zero-shot method that establishes a cross-dimensional feedback loop between 2D and 3D representations. Specifically, 2D instance masks are tracked across frames and associated with 3D superpoints, thereby integrating temporally stable 2D trajectories with spatially coherent 3D regions. This fusion enables, for the first time, the generation of globally consistent 3D instance labels in a zero-shot setting. The proposed approach significantly outperforms existing zero-shot methods across multiple benchmarks, demonstrating superior performance in terms of accuracy, temporal consistency, and scalability.
This work proposes a novel turbo inference strategy that enables bidirectional iterative refinement between detection and segmentation during inference, without requiring retraining. Challenging the conventional unidirectional “detect-then-segment” paradigm in top-down instance segmentation, the method introduces a closed-loop interaction mechanism through turbo detection and segmentation heads, coupled with cross-task feature fusion. This design dynamically leverages complementary information from both tasks while preserving the top-down architecture. Extensive experiments demonstrate significant improvements in both detection and segmentation accuracy on COCO, iFLYTEK, and Cityscapes benchmarks, achieving a favorable trade-off between performance and inference efficiency.
This work addresses the limitations of the traditional Panoptic Quality (PQ) metric, which lacks a well-defined instance matching mechanism when IoU thresholds fall below 0.5, rendering it vulnerable to challenges such as fragmentation, ambiguous boundaries, and annotation noise. The authors formulate instance matching as a constrained bipartite graph assignment problem, decoupling match confidence from both prediction and ground truth sides. They systematically define four distinct matching strategies and introduce, for the first time, a vertex-centric framework that unifies the computation of true positives, false negatives, and false positives. This approach comprehensively characterizes the space of matching strategies under low-IoU conditions and naturally extends to part-aware panoptic segmentation evaluation—particularly beneficial for biomedical image analysis. The authors further release Panoptica, an open-source evaluation toolkit supporting multi-strategy and part-level assessment, demonstrating its efficacy across multiple case studies.