Score
Designs and trains vision–language embedding models (e.g., CLIP-based) adapted to a specific target domain by applying fine-tuning, contrastive alignment, prompt/adapter, or other adaptation methods to align visual and textual representations. Builds training and evaluation procedures that preserve retrieval separability, prevent representation collapse, and enable robust zero-shot content matching between visual inputs and text.
This paper addresses the research gap in applying Contrastive Language–Image Pretraining (CLIP) to domain generalization (DG) and domain adaptation (DA). Methodologically, it establishes a unified taxonomy by systematically analyzing prevailing paradigms—including prompt optimization, backbone feature reuse, and source-available/source-free transfer—thereby elucidating CLIP’s zero-shot cross-domain transfer mechanisms and pathways to enhanced robustness. The analysis identifies three critical bottlenecks: overfitting, insufficient domain diversity, and computational inefficiency. To overcome these, the work integrates key techniques such as prompt learning, feature alignment, knowledge distillation, and domain-invariant representation learning. The contributions include both a rigorous methodological framework for CLIP-based DG/DA and actionable insights into architectural design and training strategies. Collectively, this study provides theoretical foundations and practical guidelines for developing more generalizable and deployable cross-domain vision models.
Vision-language models (VLMs) exhibit limited cross-domain generalization performance, necessitating systematic transfer strategies. Method: This work presents a comprehensive survey of VLM generalization to novel domains, proposing a modular taxonomy that unifies prompt-based, parameter-based, and feature-based transfer paradigms. It clarifies the evolutionary relationship between VLMs and multimodal large language models (MLLMs), and conducts empirical evaluation—integrating large-scale pretraining with multimodal alignment mechanisms—across mainstream benchmarks to compare methods and analyze performance. Contribution/Results: The study establishes the first technical roadmap for VLM generalization research, constructing a structured knowledge framework tailored to downstream tasks. It provides both a theoretical foundation and practical guidelines for multimodal transfer learning, advancing systematic methodology in this rapidly evolving field.
To address the dependency of downstream transfer for vision-language models on complex prompt engineering, this paper proposes CLIP-Adapter: a parameter-efficient fine-tuning method that inserts lightweight feature adapters—featuring bottleneck architectures and residual feature fusion—into either the visual or textual branch of pretrained CLIP. Unlike prevailing prompt-tuning paradigms (e.g., CoOp), CLIP-Adapter introduces feature adaptation—a novel mechanism for vision-language models—without modifying prompts or altering the original model architecture. This design preserves structural simplicity while significantly enhancing generalization across tasks. Extensive experiments demonstrate consistent superiority over state-of-the-art prompt-tuning methods on multiple image classification benchmarks. Ablation studies validate the effectiveness and cross-task transferability of each component. Overall, CLIP-Adapter establishes a new paradigm for adapting vision-language models without requiring prompt engineering.
Vision-language models like CLIP align image and text embeddings but lack semantic comparability and analogical reasoning capabilities in their embedding space, hindering vectorized reasoning about inter-image differences. Method: We propose a contrastive learning fine-tuning framework that explicitly aligns image embedding differences to LLM-generated textual difference descriptors (e.g., “thinner”, “brighter”), enabling native pairwise image difference reasoning within CLIP’s embedding space. We further introduce a novel “comparative prompting” inference paradigm to restructure the embedding geometry. Contribution/Results: After fine-tuning on synthetic difference data, our method improves average accuracy by 3.2% across attribute ranking, zero-shot classification, and retrieval tasks. It significantly enhances linear analogy preservation and directional consistency in the embedding space—providing stronger geometric foundations for downstream applications such as text-to-image generation.
This work addresses two key limitations: the weak discriminative capability of Large Vision-Language Models (LVLMs) and the insufficient language understanding and compositional reasoning of CLIP-style models. To this end, we propose the first fine-tuning framework explicitly designed to enhance discriminative ability in LVLMs. Our method jointly optimizes contrastive and autoregressive objectives, employs a multi-granularity image-text pair training strategy, and achieves parameter-efficient adaptation via synergistic soft prompting and LoRA. Crucially, it preserves the model’s generative capacity while substantially improving discriminative performance. Experiments demonstrate that our approach surpasses same-scale CLIP models on standard image-text retrieval benchmarks and achieves significant gains on compositional tasks—including VQA and visual reasoning—validating its effectiveness in strengthening deep language comprehension and structured reasoning.
To address the adaptation challenge of contrastive pre-trained vision-language models (e.g., CLIP) for few-shot classification, this paper proposes a lightweight and efficient fine-tuning method that updates only the final projection matrix of the visual encoder. Our key contributions are: (i) the first demonstration that optimizing solely this low-dimensional projection layer—without modifying the text encoder or introducing auxiliary modules—outperforms mainstream adaptation strategies; and (ii) the introduction of an L2-distance regularization between the pre-trained and fine-tuned projection matrices, which significantly enhances generalization and robustness. The method drastically reduces trainable parameters and computational overhead. It achieves state-of-the-art performance across 11 standard few-shot benchmarks and demonstrates superior results on challenging tasks including cross-domain transfer, base-to-novel class generalization, and test-time adaptation.
Soft prompt tuning often suffers from catastrophic forgetting of general-purpose knowledge in vision-language models (e.g., CLIP) under few-shot settings, leading to performance worse than zero-shot inference. Method: We propose Gradient Alignment (GA), a novel optimization mechanism that constrains prompt gradient updates to align with the direction of zero-shot predictions derived from predefined prompts—thereby explicitly preserving task-agnostic, pre-trained knowledge without requiring additional data, regularization, or architectural modifications. Contribution/Results: GA effectively mitigates overfitting and inter-class interference. It consistently outperforms state-of-the-art prompt-tuning methods across diverse transfer scenarios—including few-shot learning, domain generalization, base-to-novel class adaptation, and cross-dataset transfer—delivering substantial improvements in both generalization stability and accuracy.
To address catastrophic forgetting in fine-grained visual retrieval fine-tuning—where large-scale contrastive vision-language models degrade their general cross-modal capabilities—we propose a text-free, efficient regularization-based fine-tuning framework. Our method integrates continual learning principles with robust validation set design: (i) knowledge preservation regularization, (ii) selective fine-tuning of the visual encoder only, (iii) fine-grained hyperparameter optimization, and (iv) construction of cross-domain, reproducible validation sets—relying solely on image-side signals to maintain image–text alignment. Unlike prior approaches, our framework requires no text encoder updates or auxiliary textual annotations. Evaluated on both fine-grained and coarse-grained image–text retrieval benchmarks, it achieves state-of-the-art performance while preserving model generality and enabling domain adaptation. This work demonstrates that high-fidelity alignment can be retained through vision-only supervision, advancing efficient and scalable multimodal adaptation.
This work investigates the generalization capability of vision-language model (VLM) projection layers to unseen visual concepts—i.e., their ability to handle novel categories without explicit cross-modal alignment supervision. To this end, we introduce the first fine-grained evaluation benchmark for projection-layer generalization, built upon object detection datasets and employing a prompt-based, label-disjoint train-test split protocol. Methodologically, we integrate prompt learning, feature-space mapping, and mechanistic interpretability analysis to systematically uncover a semantic alignment mechanism in the projection layer: it functions as a class-specific key-value memory. Experiments demonstrate that the projection layer retains 79%–88% of its original performance on unseen categories—substantially outperforming standard baselines. Our findings provide both theoretical grounding and a practical paradigm for efficient, low-resource cross-modal alignment training.
Existing vision-language models (e.g., CLIP) face challenges in few-shot adaptation to fine-grained domains, including heavy reliance on prompt engineering or full-model fine-tuning, and instability or catastrophic forgetting induced by auxiliary modules. To address these issues, we propose CLIP-SVD—a parameter-efficient multimodal adaptation method based on Singular Value Decomposition (SVD). CLIP-SVD is the first to apply SVD directly to CLIP’s weight matrices, optimizing only the singular values (0.04% of total parameters), enabling cross-domain joint optimization without introducing new components and fully preserving pre-trained knowledge. Coupled with natural language analysis, it enhances interpretability. Extensive experiments across 11 natural-image and 10 biomedical datasets demonstrate that CLIP-SVD significantly outperforms state-of-the-art methods in accuracy, generalization, and adaptation efficiency.
Existing multilingual vision-language models exhibit limited cross-modal retrieval performance for low-resource languages—such as Czech, Finnish, Croatian, Hungarian, and Romanian—due to the scarcity of high-quality image–text pairs in these languages. To address this, we propose a lightweight, data-efficient language expansion method: leveraging a frozen English vision–text encoder as a semantic anchor, we train only a 1.7M-parameter cross-lingual projection module, enabling alignment without paired multilingual image–text data for the first time. Our approach operates within a contrastive learning framework, jointly optimizing frozen multilingual text and image encoders. Extensive evaluation on multiple multilingual retrieval benchmarks demonstrates substantial improvements in cross-modal retrieval performance across all five target languages. The method proves effective, generalizable across diverse low-resource settings, and deployment-friendly due to its minimal parameter overhead and reliance on frozen pretrained components.