Reprogramming Vision-Language Models via Structured Prompt Reparameterization

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of inter-class relationships and the difficulty of fine-grained discrimination in visual-language model adaptation by proposing the RVP framework. This framework models shared semantics and inter-class discrepancies through structured prompt reparameterization and introduces cross-class residual correction to reformulate visual reprogramming as a linear mapping, thereby incurring zero additional computational overhead during inference. Built upon a CLIP backbone with structured prompt aggregation and linear output aggregation, the proposed method is evaluated across eleven few-shot benchmarks and four backbone architectures. Experimental results demonstrate that RVP significantly outperforms existing approaches in both performance and efficiency.
📝 Abstract
Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggregation and do not explicitly model relationships among classes. However, fine-grained categories often exhibit highly overlapping attribute descriptions and strong inter-class correlation in the text embedding space, where discriminative cues lie in subtle low-variance components. We propose Reparameterized Inter-Class Visual Reprogramming (RVP), a structured framework that aggregates multiple text prompts within each class and applies residual correction across classes. We also show that CLIP-based visual reprogramming with input-independent linear output aggregation can be expressed as a linear mapping from frozen image embeddings to downstream logits, and use this view to design a structured reparameterization that models shared semantic components and class-specific differences. RVP uses only a single visual prompt and can be reparameterized at inference into a frozen backbone followed by a linear classifier, incurring nearly zero computational overhead. Across 11 few-shot classification benchmarks and four CLIP backbones, RVP consistently improves over prior visual reprogramming methods with comparable or better inference efficiency.
Problem

Research questions and friction points this paper is trying to address.

Visual Reprogramming
Vision-Language Models
Fine-grained Classification
Inter-class Correlation
Few-shot Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Reprogramming
Reparameterization
Inter-Class Modeling
Vision-Language Models
Few-Shot Classification
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zizhao Li
The University of Melbourne, Melbourne, Australia
Chengyi Cai
Chengyi Cai
PhD Student, the University of Melbourne
machine learningcomputer vision
M
Mohammed Yaqoob Ansari
The University of Melbourne, Melbourne, Australia
F
Feng Liu
The University of Melbourne, Melbourne, Australia
Joseph West
Joseph West
University of Melbourne
Kourosh Khoshelham
Kourosh Khoshelham
University of Melbourne
Photogrammetry3D computer visionMobile MappingSpatial InformationGeomatics