train domain-adapted vision-language models

Designs and trains vision–language embedding models (e.g., CLIP-based) adapted to a specific target domain by applying fine-tuning, contrastive alignment, prompt/adapter, or other adaptation methods to align visual and textual representations. Builds training and evaluation procedures that preserve retrieval separability, prevent representation collapse, and enable robust zero-shot content matching between visual inputs and text.

traindomain-adaptedvision-languagemodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Generalizing Vision-Language Models to Novel Domains: A Comprehensive Survey

Jun 23, 2025
XL
Xinyao Li
🏛️ University of Electronic Science and Technology of China | University of Technology Sydney | Tongji University

Vision-language models (VLMs) exhibit limited cross-domain generalization performance, necessitating systematic transfer strategies. Method: This work presents a comprehensive survey of VLM generalization to novel domains, proposing a modular taxonomy that unifies prompt-based, parameter-based, and feature-based transfer paradigms. It clarifies the evolutionary relationship between VLMs and multimodal large language models (MLLMs), and conducts empirical evaluation—integrating large-scale pretraining with multimodal alignment mechanisms—across mainstream benchmarks to compare methods and analyze performance. Contribution/Results: The study establishes the first technical roadmap for VLM generalization research, constructing a structured knowledge framework tailored to downstream tasks. It provides both a theoretical foundation and practical guidelines for multimodal transfer learning, advancing systematic methodology in this rapidly evolving field.

Comparing VLM generalization methods and benchmarks systematicallyImproving VLM performance in domain-specific generalization tasksTransferring VLM knowledge to diverse downstream applications

Must-Read Papers

Most classic and influential ideas
View more

CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Oct 09, 2021
PG
Peng Gao
🏛️ Shanghai AI Laboratory | Rutgers University | CUHK | Centre for Perceptual and Interactive Intelligence

To address the dependency of downstream transfer for vision-language models on complex prompt engineering, this paper proposes CLIP-Adapter: a parameter-efficient fine-tuning method that inserts lightweight feature adapters—featuring bottleneck architectures and residual feature fusion—into either the visual or textual branch of pretrained CLIP. Unlike prevailing prompt-tuning paradigms (e.g., CoOp), CLIP-Adapter introduces feature adaptation—a novel mechanism for vision-language models—without modifying prompts or altering the original model architecture. This design preserves structural simplicity while significantly enhancing generalization across tasks. Extensive experiments demonstrate consistent superiority over state-of-the-art prompt-tuning methods on multiple image classification benchmarks. Ablation studies validate the effectiveness and cross-task transferability of each component. Overall, CLIP-Adapter establishes a new paradigm for adapting vision-language models without requiring prompt engineering.

Enhancing CLIP performance with residual feature blendingFine-tuning visual and language branches with feature adaptersImproving vision-language models without prompt engineering

Finetuning CLIP to Reason about Pairwise Differences

Sep 15, 2024
DS
Dylan Sam
🏛️ Carnegie Mellon University | Bosch Center for AI

Vision-language models like CLIP align image and text embeddings but lack semantic comparability and analogical reasoning capabilities in their embedding space, hindering vectorized reasoning about inter-image differences. Method: We propose a contrastive learning fine-tuning framework that explicitly aligns image embedding differences to LLM-generated textual difference descriptors (e.g., “thinner”, “brighter”), enabling native pairwise image difference reasoning within CLIP’s embedding space. We further introduce a novel “comparative prompting” inference paradigm to restructure the embedding geometry. Contribution/Results: After fine-tuning on synthetic difference data, our method improves average accuracy by 3.2% across attribute ranking, zero-shot classification, and retrieval tasks. It significantly enhances linear analogy preservation and directional consistency in the embedding space—providing stronger geometric foundations for downstream applications such as text-to-image generation.

Enhance CLIP's ability to reason about image differencesEstablish geometric properties in CLIP's embedding spaceImprove zero-shot classification via comparative prompting

Discriminative Fine-tuning of LVLMs

Dec 05, 2024
YO
Yassine Ouali
🏛️ Samsung AI Cambridge | Technical University of Iasi | Queen Mary University of London

This work addresses two key limitations: the weak discriminative capability of Large Vision-Language Models (LVLMs) and the insufficient language understanding and compositional reasoning of CLIP-style models. To this end, we propose the first fine-tuning framework explicitly designed to enhance discriminative ability in LVLMs. Our method jointly optimizes contrastive and autoregressive objectives, employs a multi-granularity image-text pair training strategy, and achieves parameter-efficient adaptation via synergistic soft prompting and LoRA. Crucially, it preserves the model’s generative capacity while substantially improving discriminative performance. Experiments demonstrate that our approach surpasses same-scale CLIP models on standard image-text retrieval benchmarks and achieves significant gains on compositional tasks—including VQA and visual reasoning—validating its effectiveness in strengthening deep language comprehension and structured reasoning.

Combining contrastive and generative training for better performanceEnhancing LVLMs for discriminative vision-language tasksImproving language understanding and compositionality in vision models

Fine-Tuning CLIP's Last Visual Projector: A Few-Shot Cornucopia

Oct 07, 2024
MF
Mohammad Fahes
🏛️ Inria | Valeo.ai | Kyutai

To address the adaptation challenge of contrastive pre-trained vision-language models (e.g., CLIP) for few-shot classification, this paper proposes a lightweight and efficient fine-tuning method that updates only the final projection matrix of the visual encoder. Our key contributions are: (i) the first demonstration that optimizing solely this low-dimensional projection layer—without modifying the text encoder or introducing auxiliary modules—outperforms mainstream adaptation strategies; and (ii) the introduction of an L2-distance regularization between the pre-trained and fine-tuned projection matrices, which significantly enhances generalization and robustness. The method drastically reduces trainable parameters and computational overhead. It achieves state-of-the-art performance across 11 standard few-shot benchmarks and demonstrates superior results on challenging tasks including cross-domain transfer, base-to-novel class generalization, and test-time adaptation.

Adapting CLIP for few-shot classification without external parametersFine-tuning vision encoder's embedding projection matrix improves performanceProLIP achieves state-of-the-art in few-shot benchmarks and domain generalization

Prompt-aligned Gradient for Prompt Tuning

May 30, 2022
BZ
Beier Zhu
🏛️ Nanyang Technological University | Columbia University | Damo Academy | Alibaba Group

Soft prompt tuning often suffers from catastrophic forgetting of general-purpose knowledge in vision-language models (e.g., CLIP) under few-shot settings, leading to performance worse than zero-shot inference. Method: We propose Gradient Alignment (GA), a novel optimization mechanism that constrains prompt gradient updates to align with the direction of zero-shot predictions derived from predefined prompts—thereby explicitly preserving task-agnostic, pre-trained knowledge without requiring additional data, regularization, or architectural modifications. Contribution/Results: GA effectively mitigates overfitting and inter-class interference. It consistently outperforms state-of-the-art prompt-tuning methods across diverse transfer scenarios—including few-shot learning, domain generalization, base-to-novel class adaptation, and cross-dataset transfer—delivering substantial improvements in both generalization stability and accuracy.

Addresses improper fine-tuning undermining prompt prediction accuracyAligns gradient updates to maintain VLM generalization capabilityPrevents forgetting general knowledge during prompt tuning

Latest Papers

What's happening recently
View more

Infusing fine-grained visual knowledge to Vision-Language Models

Aug 16, 2025
NY
Nikolaos-Antonios Ypsilantis
🏛️ Czech Technical University in Prague | Google DeepMind

To address catastrophic forgetting in fine-grained visual retrieval fine-tuning—where large-scale contrastive vision-language models degrade their general cross-modal capabilities—we propose a text-free, efficient regularization-based fine-tuning framework. Our method integrates continual learning principles with robust validation set design: (i) knowledge preservation regularization, (ii) selective fine-tuning of the visual encoder only, (iii) fine-grained hyperparameter optimization, and (iv) construction of cross-domain, reproducible validation sets—relying solely on image-side signals to maintain image–text alignment. Unlike prior approaches, our framework requires no text encoder updates or auxiliary textual annotations. Evaluated on both fine-grained and coarse-grained image–text retrieval benchmarks, it achieves state-of-the-art performance while preserving model generality and enabling domain adaptation. This work demonstrates that high-fidelity alignment can be retained through vision-only supervision, advancing efficient and scalable multimodal adaptation.

Addressing catastrophic forgetting in fine-grained visual retrieval fine-tuningBalancing domain adaptation with retention of multimodal knowledgeImproving generalization across datasets without text data usage

This work investigates the generalization capability of vision-language model (VLM) projection layers to unseen visual concepts—i.e., their ability to handle novel categories without explicit cross-modal alignment supervision. To this end, we introduce the first fine-grained evaluation benchmark for projection-layer generalization, built upon object detection datasets and employing a prompt-based, label-disjoint train-test split protocol. Methodologically, we integrate prompt learning, feature-space mapping, and mechanistic interpretability analysis to systematically uncover a semantic alignment mechanism in the projection layer: it functions as a class-specific key-value memory. Experiments demonstrate that the projection layer retains 79%–88% of its original performance on unseen categories—substantially outperforming standard baselines. Our findings provide both theoretical grounding and a practical paradigm for efficient, low-resource cross-modal alignment training.

Analyzing mechanistic interpretability of projection layer as key-value memoryAssessing alignment performance on disjoint seen and unseen label setsEvaluating projection layer generalization for unseen visual concepts

Singular Value Few-shot Adaptation of Vision-Language Models

Sep 03, 2025
TK
Taha Koleilat
🏛️ Concordia University

Existing vision-language models (e.g., CLIP) face challenges in few-shot adaptation to fine-grained domains, including heavy reliance on prompt engineering or full-model fine-tuning, and instability or catastrophic forgetting induced by auxiliary modules. To address these issues, we propose CLIP-SVD—a parameter-efficient multimodal adaptation method based on Singular Value Decomposition (SVD). CLIP-SVD is the first to apply SVD directly to CLIP’s weight matrices, optimizing only the singular values (0.04% of total parameters), enabling cross-domain joint optimization without introducing new components and fully preserving pre-trained knowledge. Coupled with natural language analysis, it enhances interpretability. Extensive experiments across 11 natural-image and 10 biomedical datasets demonstrate that CLIP-SVD significantly outperforms state-of-the-art methods in accuracy, generalization, and adaptation efficiency.

Adapting vision-language models to fine-grained domains efficientlyEnhancing adaptation performance while preserving generalization abilityReducing reliance on prompt engineering and full model fine-tuning

Existing multilingual vision-language models exhibit limited cross-modal retrieval performance for low-resource languages—such as Czech, Finnish, Croatian, Hungarian, and Romanian—due to the scarcity of high-quality image–text pairs in these languages. To address this, we propose a lightweight, data-efficient language expansion method: leveraging a frozen English vision–text encoder as a semantic anchor, we train only a 1.7M-parameter cross-lingual projection module, enabling alignment without paired multilingual image–text data for the first time. Our approach operates within a contrastive learning framework, jointly optimizing frozen multilingual text and image encoders. Extensive evaluation on multiple multilingual retrieval benchmarks demonstrates substantial improvements in cross-modal retrieval performance across all five target languages. The method proves effective, generalizable across diverse low-resource settings, and deployment-friendly due to its minimal parameter overhead and reliance on frozen pretrained components.

Addressing poor multilingual retrieval performance in underrepresented languagesDeveloping parameter-efficient alignment with minimal training requirementsExtending vision-language models to low-resource languages without paired data

Hot Scholars

GP

Guansong Pang

Assistant Professor of Computer Science, Singapore Management University
Machine LearningData MiningComputer VisionAnomaly Detection
HZ

Haoyan Zhang

Shanghai Jiao Tong University
Computer Architecture
HL

Hongbin Liu

Chinese Academy of Sciences; King's College London
AI and Medical roboticsemboided AIMLLM
HL

Haoran Lai

University of Science and Technology of China
Medical Image ProcessingDeep Learning
JY

Jidong Yang

The University of Texas at Dallas; China University of Petroleum (East China)
Exploration SeismologyEarthquake SeismologyComputational Seismology