Kernel-Based Steering of CLIP with Vision-Language Model Preferences

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of transferring preferences from large vision-language models (VLMs) into CLIP’s compact embedding space. We propose ASK, a method that efficiently distills VLM preferences into CLIP without requiring teacher embeddings or additional human annotations. By leveraging positive definite target kernels and distribution anchors, ASK enables prompt-controlled visual discriminability. The approach jointly updates the encoders through the synergistic integration of low-rank adapters, visual kernel matching, and image-text distribution anchoring. Experimental results demonstrate that ASK significantly improves retrieval mean Average Precision (mAP) for ViT-B/16 from 53.8 to 75.0, while maintaining steady gains in zero-shot accuracy. These findings validate the controllability of the learned representations under specific criteria, highlighting the effectiveness of our framework in aligning lightweight models with the capabilities of larger VLMs.
πŸ“ Abstract
Large vision-language models (VLMs) can judge visual similarity, but their judgments are not directly available as compact image embeddings for efficient comparison. We study how to transfer these preferences into CLIP while retaining its image--text capabilities. We introduce ASK, a kernel-based steering method that learns from elicited pairwise judgments without accessing teacher embeddings or collecting new human similarity annotations. ASK constructs positive semidefinite target kernels within small image groups and combines visual kernel matching with an image--text distributional anchor. Low-rank adapters jointly update the visual and text encoders while regularizing predictions toward frozen CLIP. After adaptation, retrieval uses CLIP image embeddings and cosine similarity, with no VLM calls. Experiments across five image domains, four CLIP backbones, and six judges evaluate teacher agreement, retrieval, and recognition retention. For ViT-B/16, mean retrieval mAP on classes excluded from adaptation increases from 53.8 to 75.0, compared with 71.7 for DINOv2 targets with KL anchoring. Mean zero-shot accuracy with jointly adapted encoders increases from 61.8\% to 62.4\%, averaged over 12 benchmarks and the five adaptation domains. Prompting provides an additional capability: selecting which visual distinctions the student learns. Human-annotated evaluations across four datasets support this criterion-specific control.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
CLIP
visual similarity
preference transfer
image embeddings
Innovation

Methods, ideas, or system contributions that make the work stand out.

Kernel-Based Steering
Vision-Language Models
Low-Rank Adapters
Preference Transfer
Prompting Control