Seeing as Humans Do: Learning from Motion to Segment Anything Without Supervision

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of the Segment Anything Model (SAM) on massive annotations and the limited generalizability of existing unsupervised methods to static objects. To this end, we propose MoSA, a novel framework that pioneers a new paradigm for extracting transferable object priors from large-scale unlabeled videos. Specifically, MoSA trains a perceptual grouping model to learn universal object representations through multi-granularity motion pseudo-label generation, contrastive learning, and a prompt-guided architecture, thereby enabling unsupervised image segmentation. Experimental results demonstrate that our approach significantly outperforms existing unsupervised techniques on zero-shot benchmarks while achieving performance comparable to fully supervised SAM. By effectively overcoming the bottleneck of static object segmentation, MoSA establishes a promising direction for annotation-free visual perception.
📝 Abstract
The Segment Anything Model (SAM) relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling. While unsupervised methods attempt to learn object concepts from motion, they typically overfit to moving entities, lacking both multi-granularity understanding and the ability to generalize to static objects. To overcome this, we introduce Motion-Grounded Segment Anything (MoSA), a highly scalable unsupervised framework that learns a transferable objectness prior from unlabeled videos. MoSA operates in three progressive stages: (1) automatically generating multi-granularity motion pseudo-labels from large-scale video data; (2) training a Perceptual Grouping Model (PGM) via contrastive learning to internalize a generalized, appearance-driven concept of objects; and (3) transferring this learned prior into a prompt-guided architecture for segment-anything-style inference on images. Extensive zero-shot evaluations across seven challenging benchmarks (e.g., COCO and ADE20K) demonstrate that MoSA significantly outperforms existing unsupervised methods. Notably, despite using zero manual annotations, MoSA achieves segmentation performance comparable to the fully supervised SAM. Our findings reveal that harnessing large-scale unlabeled motion is a feasible and highly scalable alternative to annotation-driven segment-anything pipelines.
Problem

Research questions and friction points this paper is trying to address.

unsupervised segmentation
segment anything
motion cues
multi-granularity
zero-shot generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unsupervised Learning
Segment Anything
Motion Pseudo-labels
Perceptual Grouping
Zero-shot Segmentation
🔎 Similar Papers
No similar papers found.