🤖 AI Summary
This study addresses the high computational cost of reward-guided image editing, which typically requires repeated sampling from large models. We propose a unified theoretical framework coupled with an efficient acceleration method. The core innovation lies in decoupling the pretrained generative backbone from the inner optimization loop, substantially reducing redundant sampling through candidate reuse. Furthermore, we train a lightweight network via fast adjoint stochastic transport to perform low-cost editing, while dynamically optimizing endpoints using reward model feedback. Experimental results demonstrate that our approach comprehensively outperforms existing baselines on SD3, achieving up to 24.14× faster editing speed and realizing a significant balance between generation quality and computational efficiency.
📝 Abstract
Reward-guided image editing at test time seeks to improve a specified reward while preserving source content and visual plausibility. Many existing approaches optimize candidates through pretrained generation processes, making repeated adjustment depend on costly large-model execution and, in some cases, backbone backpropagation. We develop a theoretical framework that jointly accounts for reward, source preservation, and pretrained-prior preferences, allowing the desired output distribution to be specified separately from the dynamics used to realize it. Based on this framework, we introduce FASTER, which trains a small network for each source and objective to perform inexpensive editing, while pretrained and reward models provide feedback on candidate outputs. By reusing each candidate and its feedback across multiple small-network updates, FASTER reduces repeated sampling and supervision queries without placing the pretrained generative backbone inside the inner optimization loop. On SD3, FASTER leads all four target metrics and several validation metrics among the evaluated methods. Compared with the evaluated baseline that optimizes controls along pretrained generation trajectories, FASTER achieves editing-time speedups of up to \({6.91\times}\) on Stable Diffusion 3 and \({24.14\times}\) on Stable Diffusion 1.5.