You Only Reprogram Once: Rethinking Prolonged Training for Visual Reprogramming

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high computational cost of visual reprogramming, which typically requires hundreds of iterative optimization rounds. To this end, we propose YORO, a framework that eliminates backpropagation by constructing predictors directly from the response space of frozen models via Bayesian discriminative mapping. By integrating streaming class statistics with covariance-aware affine transformations, YORO achieves model adaptation through a single forward pass. Experimental results demonstrate that YORO improves average accuracy by 18.4%–24.4% and attains 77.2% few-shot performance on CLIP. These findings successfully validate a highly efficient β€œread-before-optimize” paradigm for model adaptation.
πŸ“ Abstract
Visual reprogramming is a parameter-efficient method for adapting pretrained models, yet its training can remain computationally expensive: even with a frozen backbone, visual prompts are often optimized through the full model for hundreds of epochs. Before changing what the pretrained model sees, we ask whether we are fully using what it already tells us. We find that modeling the full source response can already yield strong downstream predictions without prompt optimization. Motivated by this observation, we introduce You Only Reprogram Once (YORO), which constructs a downstream predictor from the frozen response space in a single forward-only traversal. Its Bayesian Discriminant Mapping (BDM) derives a covariance-aware affine mapping from streaming class statistics, requiring no backpropagation, optimizer updates, or repeated visits to the training set. When further input adaptation helps, YORO-FP optionally refines the visual prompt for 20 epochs. BDM also extends naturally to CLIP by treating attribute-prompt similarities as source responses. Across three full-data settings, YORO improves average accuracy over the strongest prior gradient-free mapping by 18.4--24.4\%. On 16-shot CLIP, it raises the four-backbone average from 71.4\% to 77.2\%. YORO-FP provides further gains on selected tasks, while validation often retains the one-pass predictor. These results suggest a different default for visual reprogramming: read out the frozen response first, and optimize the input only when needed.
Problem

Research questions and friction points this paper is trying to address.

Visual Reprogramming
Parameter-efficient adaptation
Prolonged training
Computational cost
Prompt optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Reprogramming
Bayesian Discriminant Mapping
Gradient-free Adaptation
Parameter-efficient Fine-tuning
CLIP