🤖 AI Summary
This study addresses the challenge of instruction following in image generation caused by unreliable reward models by proposing the Verifiable Visual Reward (VVR) framework. This work pioneers a programmatically verifiable image reward mechanism that constructs deterministic verifiers through synthetic scenes, enabling generalization from geometric constraints to natural language prompts. Additionally, it introduces the VVRBench benchmark and employs reinforcement learning with deterministic verifiers to post-train Stable Diffusion. Experimental results demonstrate that this approach significantly improves the accuracy of SD3.5 Medium on VVRBench from 2.8% to 28.3%, effectively enhancing complex instruction-following capabilities and alignment with human preferences.
📝 Abstract
Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, and show that training on it generalizes to natural prompts. Each VVR task is a scene of geometric objects and relations among them, from which we derive both the prompt and a deterministic verifier, so tasks can be generated in any number and at any chosen complexity. We release VVRBench, with 10,000 tasks over 32 constraint types, and VVRBench-Challenge, with 720 more complex tasks; the strongest model we evaluate---GPT-Image-2.5---solves 21.4% of VVRBench-Challenge. Using VVR scores as rewards for reinforcement learning (RLVVR) raises the accuracy of Stable Diffusion 3.5 Medium on VVRBench from 2.8% to 28.3% and demonstrates consistent easy-to-hard generalization. These gains extend to out-of-domain benchmarks, and mixing VVR into existing objectives further improves overall performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes.