Score
Design and implement training methods that learn joint image–text representations by optimizing contrastive objectives to align visual and textual embeddings or tokens. This work includes building and pairing visual and language encoders, mining or sampling hard negatives, scaling training to large image–text datasets, and measuring or improving alignment for downstream grounding and generative tasks.
This paper addresses the challenge of reducing reliance on large-scale labeled data in text-image multimodal analysis through self-supervised contrastive learning. Methodologically, it systematically reviews cross-modal positive/negative sample construction, feature-space alignment mechanisms, and unsupervised representation learning paradigms. It introduces the first taxonomy of vision-language contrastive methods based on model architecture, integrating pretraining objectives, encoder designs (e.g., CLIP, ALIGN), and similarity measurement techniques into a unified analytical framework. The work clarifies the technical evolution trajectory and identifies core bottlenecks—including computational inefficiency, sensitivity to data distribution shifts, and limited interpretability. It further proposes a modeling pathway that jointly optimizes efficiency, robustness, and explainability. The contributions provide both theoretical foundations and practical guidelines for advancing self-supervised multimodal learning, enabling more scalable, generalizable, and transparent joint representation learning across modalities.
To address insufficient text–image relational modeling and the entanglement of alignment and generation objectives caused by reliance on generic pretrained text encoders (e.g., CLIP) in text-to-image (T2I) synthesis, this paper proposes an end-to-end differentiable dual-text-embedding architecture: one branch is generation-oriented—optimized for image photorealism—while the other is alignment-oriented—designed for fine-grained text–image correspondence. The method eliminates external pretrained text encoders and jointly optimizes adversarial generation loss and contrastive alignment loss to enable cooperative learning between the two embeddings. Experiments on Oxford-102, CUB, and MS-COCO demonstrate substantial improvements over single-embedding baselines and CLIP-driven approaches, while also enabling high-fidelity text-guided image editing. The core contributions are a T2I-specific decoupled dual-embedding mechanism and a pretrained-encoder-free adaptive text representation learning framework.
To address insufficient semantic alignment in text-to-image diffusion models, this paper proposes a representation alignment (REPA)-based reconstruction training paradigm. Instead of relying solely on positive sample pairs for score matching, REPA incorporates a contrastive learning objective that jointly leverages both positive and negative sample pairs. We further design SoftREPA, a lightweight fine-tuning strategy that employs soft text tokens to achieve efficient cross-modal alignment; we theoretically prove that SoftREPA explicitly increases the mutual information between text and image representations. With fewer than 1 million additional parameters, our method significantly improves semantic consistency in both text-to-image generation and text-guided image editing tasks. Empirical evaluations demonstrate state-of-the-art performance across multiple benchmarks, outperforming existing alignment approaches.
This work challenges the prevailing paradigm that language–image alignment necessitates joint pretraining of text and image encoders (e.g., CLIP), proposing instead to freeze a pretrained large language model (LLM)—such as LLaMA or Qwen—as the fixed text encoder and perform end-to-end fine-tuning solely on a lightweight image encoder. To this end, we introduce the LIFT framework, the first systematic study demonstrating that LLM-derived text embeddings can directly and efficiently guide visual representation learning. Experiments show that our approach outperforms CLIP on compositional reasoning and long-text understanding tasks, while reducing training computational cost substantially. Our core contribution is establishing “frozen LLM + trainable image encoder” as a novel, efficient alignment paradigm—offering a simpler, more scalable alternative for multimodal representation learning.
This work addresses the challenge of learning unified cross-modal representations in image–text contrastive pretraining, where modality disconnection often hinders effective alignment. To overcome this limitation, the authors propose a lightweight fusion and multi-level alignment mechanism that operates during training: it leverages fine-grained image–text correspondences to enhance alignment and introduces a structured interaction module to mitigate early saturation in contrastive learning and improve training stability. Notably, this module is removed at inference time, preserving the efficiency of the dual-encoder architecture. Experimental results demonstrate that the proposed approach significantly outperforms strong baselines across image–text retrieval, classification, and multimodal benchmark tasks, effectively bridging the modality gap while maintaining both discriminative representation quality and inference efficiency.
This work challenges the prevailing view that image and text representations in vision-language models align only in deep layers. Inspired by DeepDream, the authors propose a synthesis method within an adapter-based architecture that extracts textual concept vectors layer by layer and optimizes corresponding images without auxiliary models or additional datasets. For the first time, this approach provides direct, concept-level evidence of cross-modal alignment starting from the very first layer: across seven network layers and hundreds of concepts, over 50% of images synthesized from the first layer already exhibit clear, salient visual features of target concepts—such as animals, activities, and seasons. This method offers an efficient and novel pathway toward enhancing the interpretability of vision-language models.
Current text-to-image generation models struggle to achieve smooth transitions between semantically similar prompts due to substantial differences in token sequences—particularly in wording, ordering, and conceptual positioning—which hinders effective image blending and continuous editing. This work proposes a Token-to-Token Alignment framework that, without modifying the underlying model, employs a two-stage strategy: first aligning the semantic structures of prompts and then aligning their token embedding representations. By reconstructing diverse prompts into a shared structural form, the method reveals that the latent continuous semantic structure within the text embedding space can be effectively leveraged through representation alignment. Consequently, linear interpolation in this aligned space yields coherent semantic transitions, significantly enhancing the quality of image semantic mixing and continuous editing.
This work proposes DREAM, a unified framework for visual representation learning and text-to-image generation. By employing a shared architecture to jointly optimize discriminative objectives (via contrastive learning) and generative objectives (via diffusion modeling), DREAM achieves synergistic gains between the two tasks. The approach incorporates a Masking Warmup training strategy and a semantic alignment decoding mechanism to enhance co-training efficacy. After training on CC12M, the model attains a 72.7% top-1 accuracy on ImageNet linear probing and a Fréchet Inception Distance (FID) of 4.25, significantly outperforming CLIP and FLUID. Moreover, it demonstrates strong generalization across diverse downstream tasks, including few-shot classification, semantic segmentation, and depth estimation.
This study investigates the reliance of text-to-image generation models on linguistic information encoded in their text encoders. To this end, the authors propose a decontextualized “bag-of-words with positional tags” embedding approach that preserves only lexical identity and word order while discarding rich semantic context, and employ it to guide a diffusion Transformer in image synthesis. Experimental results demonstrate that this simplified embedding achieves comparable performance to full contextual embeddings in both visual quality and text-image alignment. These findings reveal, for the first time, that state-of-the-art models do not critically depend on deep linguistic structures from the text encoder; instead, semantic integration is largely accomplished by the image generator itself, thereby challenging conventional assumptions about the role of text encoders in generative modeling.