🤖 AI Summary
This study addresses the reliance of the Segment Anything Model (SAM) on massive annotations and the limited generalizability of existing unsupervised methods to static objects. To this end, we propose MoSA, a novel framework that pioneers a new paradigm for extracting transferable object priors from large-scale unlabeled videos. Specifically, MoSA trains a perceptual grouping model to learn universal object representations through multi-granularity motion pseudo-label generation, contrastive learning, and a prompt-guided architecture, thereby enabling unsupervised image segmentation. Experimental results demonstrate that our approach significantly outperforms existing unsupervised techniques on zero-shot benchmarks while achieving performance comparable to fully supervised SAM. By effectively overcoming the bottleneck of static object segmentation, MoSA establishes a promising direction for annotation-free visual perception.
📝 Abstract
The Segment Anything Model (SAM) relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling. While unsupervised methods attempt to learn object concepts from motion, they typically overfit to moving entities, lacking both multi-granularity understanding and the ability to generalize to static objects. To overcome this, we introduce Motion-Grounded Segment Anything (MoSA), a highly scalable unsupervised framework that learns a transferable objectness prior from unlabeled videos. MoSA operates in three progressive stages: (1) automatically generating multi-granularity motion pseudo-labels from large-scale video data; (2) training a Perceptual Grouping Model (PGM) via contrastive learning to internalize a generalized, appearance-driven concept of objects; and (3) transferring this learned prior into a prompt-guided architecture for segment-anything-style inference on images. Extensive zero-shot evaluations across seven challenging benchmarks (e.g., COCO and ADE20K) demonstrate that MoSA significantly outperforms existing unsupervised methods. Notably, despite using zero manual annotations, MoSA achieves segmentation performance comparable to the fully supervised SAM. Our findings reveal that harnessing large-scale unlabeled motion is a feasible and highly scalable alternative to annotation-driven segment-anything pipelines.