ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of prompt learning in vision-language models under few-shot and distribution-shift scenarios, specifically weak cross-modal interaction and susceptibility to overfitting. To this end, we propose a probabilistic cross-attention prompting framework that introduces a stacked bidirectional multi-head cross-attention mechanism to enhance cross-modal feature refinement. Prompt tokens are modeled as Gaussian distributions parameterized by learnable means and variances, while lightweight KL divergence and L2 regularization are incorporated to mitigate overfitting, with the CLIP backbone kept frozen throughout training. Extensive experiments demonstrate that the proposed framework achieves superior performance across 11 datasets on few-shot generalization, cross-dataset transfer, and domain generalization tasks, effectively improving model robustness and stability.
📝 Abstract
Pre-trained vision-language models such as CLIP can recognize new categories via prompting, but they often struggle when labeled data are scarce or the test distribution shifts. Prompt learning adapts only a small set of parameters while keeping the backbone frozen, yet many existing multimodal prompt learners couple the visual and textual branches weakly and can be brittle in low-shot regimes. We propose ProCAP, a probabilistic cross-attentive prompt learning framework that improves cross-modal interaction and training stability without updating any CLIP weights: it learns both visual and textual prompt tokens and links them through stacked bidirectional multi-head cross-attention so the two branches refine each other across prompt depth. To reduce overfitting under limited supervision, we parameterize prompt tokens with Gaussian means and variances and regularize them with lightweight KL and L2 penalties, and we further add a compact symmetric InfoNCE head that aligns cross-attended image features with class-level text representations in a shared low-dimensional space. Across few-shot base-to-novel generalization on 11 datasets, cross-dataset transfer, and domain generalization on ImageNet shift benchmarks, ProCAP achieves strong aggregate base-to-novel performance and competitive transfer performance while keeping the CLIP backbone unchanged.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Prompt Learning
Few-shot Learning
Domain Generalization
Cross-modal Interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Probabilistic Prompt Learning
Cross-Attention
Vision-Language Models
Few-Shot Generalization
InfoNCE Alignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hiwa Azeez Abbas
University of Kurdistan, Sanandaj, Iran
F
Fatemeh Daneshfar
Department of Computer Engineering, University of Kurdistan, Sanandaj, Iran
Moloud Abdar
Moloud Abdar
Senior Data Scientist, The University of Queensland, Australia
Machine LearningDeep LearningComputer VisionVision-Language ModelsSentiment Analysis