Score
Designs and implements unsupervised segmentation systems that produce pixel- or region-level partitions of image and video data without using ground-truth labels, covering binary/object masks, motion-based segments, and clustering-derived region groupings. Builds algorithms and evaluation analyses for generating coarse or initial masks for refinement, scaling segmentation to high-resolution inputs, and ensuring temporal or spatial consistency across frames and regions.
Self-supervised learning (SSL) for image segmentation often relies heavily on large-scale annotated data, hindering its practical deployment. This paper presents a comprehensive survey of over 150 state-of-the-art SSL-based segmentation works published between 2018 and 2024. Methodologically, it introduces the first three-dimensional taxonomy—spanning pretraining paradigms (e.g., contrastive learning, generative modeling, redundancy reduction, geometric transformation), downstream tasks (semantic, instance, and medical segmentation), and benchmark datasets—and establishes a reproducible, unified evaluation framework to uncover methodological commonalities and performance limits. Furthermore, it constructs a knowledge graph integrating methodologies, datasets, benchmark results, and critical limitations analysis. The contributions significantly lower the entry barrier for researchers, provide theoretical foundations for standardization, and offer practical guidelines for real-world adoption of SSL in segmentation.
This work introduces the first unsupervised video panoptic segmentation task and proposes VideoCUPS, a method that detects, segments, and tracks all objects in videos without any human annotations. Leveraging depth, motion, and visual cues from scene-centric videos, VideoCUPS generates temporally consistent pseudo-labels and employs a novel Video DropLoss for optimized training. The study also establishes a comprehensive evaluation protocol and baseline models. Experimental results demonstrate that VideoCUPS significantly outperforms various baselines derived from state-of-the-art image or instance segmentation models, highlighting its strong performance in label-efficient learning.
This work addresses the high cost of pixel-level annotations in video semantic segmentation by proposing an efficient learning paradigm that leverages unlabeled video frames alongside coarse-grained labels. The approach utilizes the Segment Anything Model (SAM/SAM 2) to automatically generate and refine segmentation masks. Systematic evaluation demonstrates, for the first time, that this method reduces human annotation effort by approximately 33% while maintaining comparable performance to fully supervised baselines. Furthermore, the study reveals that inter-frame diversity exerts a substantially greater influence on model performance than the sheer number of frames, underscoring the critical role of data diversity in weakly supervised video segmentation.
Unsupervised video object segmentation (UVOS) suffers from heavy reliance on frame-wise mask annotations and limited generalization. To address this, we propose the first mask-free UVOS paradigm: for the first time, we adapt the Segment Anything Model (SAM) to the video domain, enabling temporal-consistent segmentation using only learnable bounding boxes as prompts—without any mask supervision. Our key contributions are: (1) STD-Net, a tracker featuring spatio-temporal decoupled deformable attention, significantly enhancing robustness and cross-frame consistency of box prompts under complex scenes; and (2) a prompt-driven video propagation framework coupled with unsupervised temporal feature alignment. Experiments demonstrate state-of-the-art performance on DAVIS2017-Unsupervised and YouTube-VIS 2019/2021, surpassing mainstream supervised methods despite zero mask supervision—achieving a J&F score of 68.3%. Moreover, our method exhibits strong generalization to weakly annotated data.
While the Segment Anything Model (SAM) exhibits strong generalization capability, its substantial computational overhead hinders deployment on resource-constrained edge devices. This work presents a systematic survey of efficient SAM variants tailored for edge deployment. We introduce the first unified evaluation framework spanning diverse hardware platforms—including CPU, GPU, and Edge TPU—and conduct joint accuracy–latency–memory benchmarking on COCO and SA-1B. Our analysis categorizes acceleration techniques along six technical axes: model pruning, knowledge distillation, lightweight attention mechanisms, quantization, module substitution, and hardware-aware compilation—characterizing their Pareto-optimal trade-offs. The core contributions are: (1) an open-source, fully reproducible edge-SAM benchmark; and (2) empirical insights into the applicability domains and fundamental accuracy-efficiency trade-offs of each acceleration strategy—providing both theoretical foundations and practical guidelines for designing lightweight vision foundation models.
In unsupervised video object segmentation, models often suffer from unstable predictions due to over-reliance on motion cues such as optical flow. To address this, we propose the “Motion-as-Option” mechanism, which decouples motion information into an optional module: during training, optical flow inputs to the motion encoder are randomly replaced with RGB frames, and an adaptive output selection algorithm dynamically fuses predictions from parallel motion and appearance pathways. This is the first approach to enable non-mandatory modeling of motion cues, integrating motion-appearance collaborative representation learning with stochastic input masking. Our method achieves state-of-the-art performance on DAVIS and FBMS benchmarks, demonstrating significantly improved robustness against anomalous motion disturbances and yielding a 23% gain in prediction stability.
To address the heavy reliance of instance segmentation on costly, labor-intensive manual annotations, this paper proposes a fully unsupervised instance segmentation framework. Methodologically, it introduces the first integration of superpixels—generated via MultiCut and low-level features—with self-supervised visual representations, and designs a superpixel-guided mask loss with dual hard and soft branches. Furthermore, an adaptive-weighted self-training mechanism is incorporated to enable pseudo-label quality-driven iterative optimization. The core contributions are: (1) joint modeling of superpixels and self-supervised features; (2) a two-stage learnable mask loss function; and (3) an adaptive self-training strategy. Evaluated on standard benchmarks, the proposed method achieves state-of-the-art performance in both unsupervised instance segmentation and unsupervised object detection, outperforming all existing approaches.
This work addresses the challenge of insufficient image segmentation accuracy under complex conditions such as severe illumination non-uniformity by proposing a two-stage clustering method. In the first stage, superpixels are generated via linear least-squares assignment; in the second stage, these superpixels are greedily merged into semantic regions based on the squared 2-Wasserstein distance between their empirical distributions. The key innovation lies in the novel introduction of discrete optimal transport into the superpixel merging phase, replacing conventional mean-color metrics with the squared 2-Wasserstein distance to achieve mathematical consistency across both clustering levels. Experimental results demonstrate that the proposed approach significantly improves segmentation accuracy on challenging images while maintaining high computational efficiency.
Existing unsupervised video instance segmentation methods heavily rely on synthetic videos—e.g., generated by translating/scaling ImageNet images—and thus fail to capture complex real-world motion dynamics, including viewpoint changes, multi-object interactions, and camera motion. This work presents the first end-to-end unsupervised framework trained exclusively on real-world videos. We introduce a KeyMask selection mechanism to automatically identify high-quality, sparse keyframe masks; propose a sparse-to-dense distillation framework that integrates deep motion priors with an implicit mask propagation network to generate temporally consistent dense masks; and design a novel Temporal DropLoss to improve robustness against inter-frame perturbations and motion discontinuities. Our method achieves significant improvements over state-of-the-art approaches across multiple benchmarks, with substantial gains in both segmentation accuracy and temporal stability.
To address unsupervised segmentation of large-scale unlabeled images (e.g., advertisements, social media content), this paper proposes CLASP—a lightweight, training-free, annotation-free, and hyperparameter-free framework. Methodologically, CLASP extracts local features using DINO-ViT, constructs a similarity matrix for adaptive spectral clustering (automatically determining the optimal number of clusters), and refines segment boundaries via feature sharpening and DenseCRF post-processing. Its core contribution is an end-to-end segmentation pipeline that eliminates model training, manual labeling, and hyperparameter tuning, thereby significantly enhancing reproducibility and deployment efficiency. Evaluated on COCO-Stuff and ADE20K, CLASP achieves mIoU and pixel accuracy competitive with state-of-the-art unsupervised methods. These results validate its practical utility and generalizability in real-world applications such as brand safety monitoring and creative asset management.
SAM-family models lack explicit, continuous control over segmentation granularity; users rely on manual prompt engineering or post-hoc mask filtering—processes that are ambiguous and poorly generalizable. Method: We propose the first annotation-free, arbitrary-granularity image segmentation framework, extending SAM-2 via self-supervised learning. Our approach introduces a lightweight (0.02% parameter increase) granularity-aware module and a novel granularity-control embedding mechanism, coupled with an unsupervised partitioning strategy trained on only 6K unlabeled images to enable fine-grained, continuous granularity modulation. The method supports interactive, full-image, and video segmentation. Results: Evaluated across 11 benchmarks, our method achieves substantial improvements: Number of Clicks to reach 90% IoU (NoC90) decreases from 5.69 to 4.75; 1−IoU improves to 73.1; and Average Recall at 1000 proposals (AR1000) rises to 68.3.