🤖 AI Summary
This study addresses the absence of inter-class relationships and the difficulty of fine-grained discrimination in visual-language model adaptation by proposing the RVP framework. This framework models shared semantics and inter-class discrepancies through structured prompt reparameterization and introduces cross-class residual correction to reformulate visual reprogramming as a linear mapping, thereby incurring zero additional computational overhead during inference. Built upon a CLIP backbone with structured prompt aggregation and linear output aggregation, the proposed method is evaluated across eleven few-shot benchmarks and four backbone architectures. Experimental results demonstrate that RVP significantly outperforms existing approaches in both performance and efficiency.
📝 Abstract
Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggregation and do not explicitly model relationships among classes. However, fine-grained categories often exhibit highly overlapping attribute descriptions and strong inter-class correlation in the text embedding space, where discriminative cues lie in subtle low-variance components. We propose Reparameterized Inter-Class Visual Reprogramming (RVP), a structured framework that aggregates multiple text prompts within each class and applies residual correction across classes. We also show that CLIP-based visual reprogramming with input-independent linear output aggregation can be expressed as a linear mapping from frozen image embeddings to downstream logits, and use this view to design a structured reparameterization that models shared semantic components and class-specific differences. RVP uses only a single visual prompt and can be reparameterized at inference into a frozen backbone followed by a linear classifier, incurring nearly zero computational overhead. Across 11 few-shot classification benchmarks and four CLIP backbones, RVP consistently improves over prior visual reprogramming methods with comparable or better inference efficiency.