π€ AI Summary
To address unsupervised segmentation of large-scale unlabeled images (e.g., advertisements, social media content), this paper proposes CLASPβa lightweight, training-free, annotation-free, and hyperparameter-free framework. Methodologically, CLASP extracts local features using DINO-ViT, constructs a similarity matrix for adaptive spectral clustering (automatically determining the optimal number of clusters), and refines segment boundaries via feature sharpening and DenseCRF post-processing. Its core contribution is an end-to-end segmentation pipeline that eliminates model training, manual labeling, and hyperparameter tuning, thereby significantly enhancing reproducibility and deployment efficiency. Evaluated on COCO-Stuff and ADE20K, CLASP achieves mIoU and pixel accuracy competitive with state-of-the-art unsupervised methods. These results validate its practical utility and generalizability in real-world applications such as brand safety monitoring and creative asset management.
π Abstract
We introduce CLASP (Clustering via Adaptive Spectral Processing), a lightweight framework for unsupervised image segmentation that operates without any labeled data or finetuning. CLASP first extracts per patch features using a self supervised ViT encoder (DINO); then, it builds an affinity matrix and applies spectral clustering. To avoid manual tuning, we select the segment count automatically with a eigengap silhouette search, and we sharpen the boundaries with a fully connected DenseCRF. Despite its simplicity and training free nature, CLASP attains competitive mIoU and pixel accuracy on COCO Stuff and ADE20K, matching recent unsupervised baselines. The zero training design makes CLASP a strong, easily reproducible baseline for large unannotated corpora especially common in digital advertising and marketing workflows such as brand safety screening, creative asset curation, and social media content moderation