VLM4Cluster: Benchmarking Deep Clustering In the Era of Vision-Language Pre-training

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inconsistent evaluation criteria, limited datasets, and narrow assessment dimensions in image clustering during the vision-language pretraining era. To this end, we construct a comprehensive benchmark encompassing 17 methods and 20 datasets. By unifying experimental settings and incorporating analyses of adversarial perturbations, distribution shifts, and computational efficiency, this work systematically quantifies, for the first time, the genuine gains and limitations of language-assisted clustering in complex scenarios. Our findings reveal that linguistic priors substantially enhance clustering performance and generalization in semantically dense tasks; however, such benefits diminish in large-scale, fine-grained settings.
📝 Abstract
Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics to complement visual representations. Despite the rapid proliferation of LaIC methods, it remains unclear how much LaIC has actually advanced image clustering, as existing studies generally suffer from major limitations, including inconsistent experimental settings, inadequate dataset selection, and limited evaluation dimensions. To address this gap, we introduce VLM4Cluster, a comprehensive benchmark for image clustering in the era of pre-trained vision-language models (VLMs). VLM4Cluster implements 17 representative methods spanning classical, deep, and language-assisted image clustering, and evaluates them on 20 datasets covering classical, challenging, fine-grained, large-scale, and out-of-distribution settings. Beyond effectiveness, VLM4Cluster systematically investigates image clustering along three complementary dimensions: robustness to adversarial perturbations, generalization under distribution shifts, and computational efficiency. Our study shows that LaIC substantially advances the clustering performance frontier on many semantically demanding benchmarks, generally exhibits stronger generalization under distribution shifts, and achieves a more favorable effectiveness-efficiency trade-off. However, its gains become less consistent on large-scale and fine-grained datasets, while language assistance does not systematically reduce sensitivity to adversarial perturbations. VLM4Cluster is released at https://github.com/YuanweiHuu/VLM4Cluster.
Problem

Research questions and friction points this paper is trying to address.

image clustering
vision-language pre-training
benchmarking
language-assisted image clustering
evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Pre-training
Image Clustering
Benchmark
Language-assisted Image Clustering
Robustness
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yuanwei Hu
University College London
B
Bo Peng
The University of Queensland
Y
Yuheng Jia
Southeast University
Xinting Hu
Xinting Hu
Max Planck Institute for Informatics
Multimodal ReasoningContinual LearningSemi-Supervised Learning
Yadan Luo
Yadan Luo
ARC DECRA and Senior Lecturer, University of Queensland
Generalization3D VisionAutonomous Driving
W
Wenjie Zhu
Auckland University of Technology