Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision

📅 2025-08-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the generalization capability of vision-language model (VLM) projection layers to unseen visual concepts—i.e., their ability to handle novel categories without explicit cross-modal alignment supervision. To this end, we introduce the first fine-grained evaluation benchmark for projection-layer generalization, built upon object detection datasets and employing a prompt-based, label-disjoint train-test split protocol. Methodologically, we integrate prompt learning, feature-space mapping, and mechanistic interpretability analysis to systematically uncover a semantic alignment mechanism in the projection layer: it functions as a class-specific key-value memory. Experiments demonstrate that the projection layer retains 79%–88% of its original performance on unseen categories—substantially outperforming standard baselines. Our findings provide both theoretical grounding and a practical paradigm for efficient, low-resource cross-modal alignment training.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
Vision-Language Models (VLMs) combine a vision encoder and a large language model (LLM) through alignment training, showing strong performance on multimodal tasks. A central component in this architecture is the projection layer, which maps visual features into the LLM's embedding space. Despite its importance, its ability to generalize to unseen visual concepts has not been systematically evaluated. To address this, we propose a benchmark for evaluating projection-layer generalization. We adapt object detection datasets (rich in fine-grained annotations) into a prompting format and design train/test splits with disjoint label sets, enabling precise control over seen and unseen concept separation. Experimental results show that the projection layer retains about 79 to 88 percent of the performance on unseen classes compared to seen ones across various settings, suggesting a non-trivial level of generalization even without explicit alignment supervision on those concepts. We further analyze this behavior through a mechanistic interpretability lens. Our findings indicate that the feed-forward network in the projection layer functions like a key-value memory, processing seen and unseen tokens in similar ways. This study introduces a new evaluation framework for alignment generalization and highlights the potential for efficient VLM training with limited aligned data.
Problem

Research questions and friction points this paper is trying to address.

Evaluating projection layer generalization for unseen visual concepts
Assessing alignment performance on disjoint seen and unseen label sets
Analyzing mechanistic interpretability of projection layer as key-value memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Projection layer generalization benchmark using disjoint labels
Mechanistic interpretability reveals key-value memory functionality
Evaluation framework enables efficient VLM training with limited data
💼 Related Jobs
No related jobs found.
R
Raehyuk Jung
Computer Vision and Machine Learning (CVML), KAIST, Seocho-gu, Seoul, South Korea
S
Seungjun Yu
Computer Vision and Machine Learning (CVML), KAIST, Seocho-gu, Seoul, South Korea
Hyunjung Shim
Hyunjung Shim
Associate Professor, KAIST
Computer visionmachine learning