Unlocking Fine-Grained Perception in CLIP via Structurally-Aware Latent Masked Modeling

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited fine-grained perception capability of CLIP, which constrains its visual potential in dense prediction and multimodal large language models. To this end, we propose SALM, an unsupervised framework that synergizes local and global alignment via a dual-path design to enhance CLIP representations without requiring image-text pairs. Specifically, SALM introduces structure-aware latent mask modeling, incorporating explicit calibration and implicit aggregation mechanisms to inject geometric priors and recover semantic details. We further derive an efficient self-distillation paradigm, termed SALM-Self. Extensive experiments demonstrate that our approach substantially improves dense prediction performance and zero-shot accuracy, effectively bolstering the fine-grained understanding capabilities of multimodal large models.
📝 Abstract
Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense prediction tasks and bottlenecks the visual potential of Multimodal Large Language Models (MLLMs). Existing research has attempted to enhance CLIP's visual representations by incorporating geometric priors from vision-centric models. However, these strategies often struggle to achieve deep alignment for both local spatial structures and global semantics, potentially even distorting the original image-text space. To address these limitations, we propose SALM, an unsupervised embedding alignment framework based on structurally-aware latent mask modeling. SALM effectively synergizes local and global alignment via a dual-path design combining explicit and implicit mechanisms, without requiring any image-text pairs. First, we introduce a dual-matrix alignment strategy that explicitly calibrates intra-sample spatial correlations and activation intensities, thereby effectively injecting local geometric priors. Based on this, we further design a latent mask modeling mechanism to guide CLIP to restore the missing semantic details of the target model, thereby implicitly aggregating fine-grained structures into the global semantic space. Furthermore, driven by the empirical observations that CLIP's shallow features inherently possess strong spatial observational capabilities, we naturally extend SALM to a highly efficient self-distillation paradigm, SALM-Self. This unlocks CLIP's intrinsic fine-grained potential without relying on any external models. Extensive experiments demonstrate that SALM not only significantly improves performance in dense prediction tasks but also boosts CLIP's zero-shot accuracy, effectively enhancing the fine-grained understanding capabilities of MLLMs. Project page at https://qzfm.github.io/salm_project_page/.
Problem

Research questions and friction points this paper is trying to address.

Fine-Grained Perception
Vision-Language Models
Dense Prediction
Multimodal Large Language Models
Semantic Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structurally-Aware Latent Masked Modeling
Dual-Matrix Alignment
Unsupervised Embedding Alignment
Self-Distillation
Fine-Grained Perception
💼 Related Jobs
No related jobs found.
J
Juntong Li
School of Software Engineering, South China University of Technology
Lingwei Dang
Lingwei Dang
South China University of Technology
Computer visionAIGCMultimodal learningEmbodied AI
H
Haomin Wu
School of Software Engineering, South China University of Technology
Z
Ziyan Qiu
School of Software Engineering, South China University of Technology
Qingxin Xiao
Qingxin Xiao
South China University of Technology
MLLMTask oriented dialogue
Qingyao Wu
Qingyao Wu
School of Software Engineering, South China University of Technology
Computer VisionMachine Learning