What Makes Synthetic Hard Negatives Work in Vision-Language Pretraining?

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue in vision-language pretraining where cross-modal synthetic negatives frequently contain false positives, leading to transfer failures. To mitigate this, we propose SNAP, a method that introduces a purely intra-modal hard negative generation strategy. By leveraging geometric analysis of the representation space, SNAP circumvents failure modes and incorporates a fixed temperature scaling mechanism to optimize training, thereby completely eliminating positive sample interference. The approach is model-agnostic and computationally lightweight, seamlessly integrating with CLIP and FLIP architectures. Experimental results demonstrate that SNAP yields consistent performance improvements across zero-shot retrieval, classification, and linear probing tasks, while incurring less than 10% additional training overhead.
📝 Abstract
Synthetic hard negatives generated in the representation space have proven effective for unimodal self-supervised learning, but transferring this idea to vision-language pretraining is not straightforward. We analyze six representation-space synthesis strategies and identify two failure modes in their transfer to vision-language pretraining: cross-modal constructions that produce overly easy negatives or pull them toward the query, and intra-modal constructions that incorporate the matched positive. We also observe logit-scale saturation when training with synthetic hard negatives and a learnable temperature, and find that fixing the temperature improves downstream performance. Using this geometric analysis we propose SNAP, which generates intra-modal hard negatives that never involve the positive from either modality, avoiding both failure modes entirely. SNAP is model-agnostic, requires no external generative models, and adds less than 10% training time overhead. Evaluated on top of CLIP and FLIP across multiple architectures and datasets, SNAP delivers consistent improvements on zero-shot retrieval, zero-shot classification, and linear probe evaluation.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Pretraining
Synthetic Hard Negatives
Representation Space
Contrastive Learning
Logit-Scale Saturation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic Hard Negatives
Vision-Language Pretraining
Representation Space
Geometric Analysis
SNAP