CLASP: Adaptive Spectral Clustering for Unsupervised Per-Image Segmentation

πŸ“… 2025-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
To address unsupervised segmentation of large-scale unlabeled images (e.g., advertisements, social media content), this paper proposes CLASPβ€”a lightweight, training-free, annotation-free, and hyperparameter-free framework. Methodologically, CLASP extracts local features using DINO-ViT, constructs a similarity matrix for adaptive spectral clustering (automatically determining the optimal number of clusters), and refines segment boundaries via feature sharpening and DenseCRF post-processing. Its core contribution is an end-to-end segmentation pipeline that eliminates model training, manual labeling, and hyperparameter tuning, thereby significantly enhancing reproducibility and deployment efficiency. Evaluated on COCO-Stuff and ADE20K, CLASP achieves mIoU and pixel accuracy competitive with state-of-the-art unsupervised methods. These results validate its practical utility and generalizability in real-world applications such as brand safety monitoring and creative asset management.

Technology Category

Computer Vision: SegmentationMachine Learning: Unsupervised & Self-Supervised LearningNatural Language Processing: Summarization

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphs
πŸ“ Abstract
We introduce CLASP (Clustering via Adaptive Spectral Processing), a lightweight framework for unsupervised image segmentation that operates without any labeled data or finetuning. CLASP first extracts per patch features using a self supervised ViT encoder (DINO); then, it builds an affinity matrix and applies spectral clustering. To avoid manual tuning, we select the segment count automatically with a eigengap silhouette search, and we sharpen the boundaries with a fully connected DenseCRF. Despite its simplicity and training free nature, CLASP attains competitive mIoU and pixel accuracy on COCO Stuff and ADE20K, matching recent unsupervised baselines. The zero training design makes CLASP a strong, easily reproducible baseline for large unannotated corpora especially common in digital advertising and marketing workflows such as brand safety screening, creative asset curation, and social media content moderation
Problem

Research questions and friction points this paper is trying to address.

Automates unsupervised image segmentation without labeled data or training
Automatically selects segment count and sharpens boundaries adaptively
Provides reproducible baseline for unannotated image analysis workflows
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses self-supervised ViT encoder for patch features
Automates segment count via eigengap silhouette search
Sharpens boundaries with fully connected DenseCRF
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
M
Max Curie
Integral Ad Science, New York, USA
Paulo da Costa
Paulo da Costa
Integral Ad Science, New York, USA