Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the catastrophic forgetting in CLIP post-training caused by improper temperature parameters in contrastive loss, proposing the ComCLIP framework. The method freezes the text encoder and fine-tunes only the vision encoder for a single epoch. By revealing and employing a small-temperature InfoNCE loss to mitigate forgetting, it further optimizes representations through MSE anchoring and DINOv2 relational distillation. This approach requires no architectural modifications and incurs zero additional inference overhead. Experiments demonstrate that ComCLIP significantly improves linear probing transfer performance (48.99 vs. 42.28), enhances MMVP scores, and preserves compatibility with downstream vision-language models, thereby providing a lightweight and efficient post-training paradigm for vision encoders.
📝 Abstract
CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA. Post-training offers a lightweight route to refine CLIP, but recent work argues that the standard contrastive loss is unsuitable for post-training due to catastrophic forgetting under small batches, motivating designs that abandon the contrastive objective in favor of distillation. We revisit this premise and find that, for the InfoNCE objective, the reported forgetting is driven primarily not by insufficient negatives but by an inappropriate magnitude of the contrastive temperature $τ$: with $τ$ set sufficiently small, contrastive post-training improves rather than degrades the pretrained CLIP, which we explain through the temperature dependence of the InfoNCE gradient. Building on this finding, we propose \textbf{ComCLIP}, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2. Over multiple seeds, ComCLIP matches the self-distillation baseline CLIP-Refine on zero-shot classification while significantly improving the transferability of visual features, measured by linear probing ($48.99$ vs.\ $42.28$ on ViT-B/16), and on ViT-L/14 it also improves MMVP over CLIP-Refine ($24.20$ vs.\ $19.01$); CLIP-Refine remains stronger on image-text retrieval. Used as a drop-in vision encoder for LLaVA-1.5-7B without re-aligning the projector or LLM, ComCLIP yields no net change across $8$ VLM benchmarks, i.e., the refinement does not break downstream compatibility. Code and models are available at https://github.com/showstarpro/ComCLIP.git.
Problem

Research questions and friction points this paper is trying to address.

CLIP post-training
contrastive loss
catastrophic forgetting
InfoNCE temperature
vision-language model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contrastive Loss
Temperature Scaling
Post-training
Drop-in Replacement
Relational Distillation
🔎 Similar Papers
2024-04-30International Conference on Machine LearningCitations: 13
💼 Related Jobs
No related jobs found.
Z
Zidan Wang
Huazhong University of Science and Technology
Yaqian Li
Yaqian Li
Li Auto
computer vision
X
Xiaokai Zhang
Huazhong University of Science and Technology
K
Kaiwen Long
Li Auto Inc.
Kun He
Kun He
Professor, Huazhong University of Science and Technology
AI SecurityGraph data miningOptimizationDeep learningAI4Sci
H
Hanpeng Liu
Huazhong University of Science and Technology, Li Auto Inc.