🤖 AI Summary
This study addresses the limited fine-grained perception capability of CLIP, which constrains its visual potential in dense prediction and multimodal large language models. To this end, we propose SALM, an unsupervised framework that synergizes local and global alignment via a dual-path design to enhance CLIP representations without requiring image-text pairs. Specifically, SALM introduces structure-aware latent mask modeling, incorporating explicit calibration and implicit aggregation mechanisms to inject geometric priors and recover semantic details. We further derive an efficient self-distillation paradigm, termed SALM-Self. Extensive experiments demonstrate that our approach substantially improves dense prediction performance and zero-shot accuracy, effectively bolstering the fine-grained understanding capabilities of multimodal large models.
📝 Abstract
Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense prediction tasks and bottlenecks the visual potential of Multimodal Large Language Models (MLLMs). Existing research has attempted to enhance CLIP's visual representations by incorporating geometric priors from vision-centric models. However, these strategies often struggle to achieve deep alignment for both local spatial structures and global semantics, potentially even distorting the original image-text space. To address these limitations, we propose SALM, an unsupervised embedding alignment framework based on structurally-aware latent mask modeling. SALM effectively synergizes local and global alignment via a dual-path design combining explicit and implicit mechanisms, without requiring any image-text pairs. First, we introduce a dual-matrix alignment strategy that explicitly calibrates intra-sample spatial correlations and activation intensities, thereby effectively injecting local geometric priors. Based on this, we further design a latent mask modeling mechanism to guide CLIP to restore the missing semantic details of the target model, thereby implicitly aggregating fine-grained structures into the global semantic space. Furthermore, driven by the empirical observations that CLIP's shallow features inherently possess strong spatial observational capabilities, we naturally extend SALM to a highly efficient self-distillation paradigm, SALM-Self. This unlocks CLIP's intrinsic fine-grained potential without relying on any external models. Extensive experiments demonstrate that SALM not only significantly improves performance in dense prediction tasks but also boosts CLIP's zero-shot accuracy, effectively enhancing the fine-grained understanding capabilities of MLLMs. Project page at https://qzfm.github.io/salm_project_page/.