One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inconsistent impact of the modality gap in vision-language models on downstream tasks by proposing a unified geometric interpretation framework. Building upon CLIP and SigLIP encoders, this work employs rank-one decomposition and geometric analysis of similarity scores to elucidate the geometric nature of the modality gap and its underlying mechanisms across different readout strategies. Our findings reveal that the dominant direction captures between 94.4% and 99.9% of the separation norm, thereby providing a unified explanation for the differential effects of gap modification on tasks such as classification and retrieval. Consequently, this framework offers a solid theoretical foundation for selecting appropriate gap intervention strategies.
📝 Abstract
Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval. In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one. A decomposition of the similarity score then identifies three task-specific roles. In zero-shot classification, query-side fixed gap-offset subtraction is exactly equivalent to an additive class bias. In standard cross-modal retrieval, projecting out the gap direction and renormalising residuals discards candidate-specific norm information, inducing a multiplicative ranking distortion; a geometry-derived exponent tracks the grid-search optimum (Spearman rho = 0.93) and restores performance in some settings, although the gains transfer unevenly. In mixed-modal retrieval, the gap direction sorts candidates by modality; its removal can improve cross-modal ranking, unlike random or non-gap controls. Residual semantic structure after removal defines the limits of the rank-one account. Together, these results explain why gap modification can improve, degrade, or restore performance across downstream settings. By clarifying when and why gap modification changes model behavior, this account provides a principled basis for selecting gap interventions in similarity-based vision-language systems across evaluated downstream tasks.
Problem

Research questions and friction points this paper is trying to address.

modality gap
vision-language models
contrastive learning
downstream tasks
representation geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

modality gap
vision-language models
geometric decomposition
rank-one structure
cross-modal retrieval
💼 Related Jobs
No related jobs found.