🤖 AI Summary
This study addresses the limitation of existing multi-subject image generation methods, which merely verify subject presence without ensuring that attributes, actions, and relationships are correctly bound to their reference subjects. To overcome this, the authors propose Visual Jev, a reference-binding-based visual reward mechanism. Leveraging the Qwen3.5-4B verifier and the MICo-150K dataset, binary supervision signals are constructed offline through fixed question-answering pairs. The probability of affirmative responses from the language model is then utilized as a reinforcement learning reward within the GRPO framework, effectively translating visual judgments into training signals. Evaluated on an 897-task subset, the approach improves the GPT-5.4 composite score from 41.78 to 52.50. While the work provides a comprehensive implementation and evaluation framework, the statistical significance of these results warrants further verification.
📝 Abstract
Multi-subject image generation requires rewards that verify whether requested attributes, actions, and relations hold for the specified reference subjects. Subject presence alone does not establish that the correct subjects participate in a requested interaction. We present reference-bound Visual Jev rewards that turn these visual decisions into generator training signals. Each subject-related question receives a positive label only when the requested condition and the relevant reference identities hold jointly. We construct fixed questions offline, train a Qwen3.5-4B verifier with binary supervision, and directly read Yes probabilities from its language-model head. Their mean supplies a GRPO reward while retaining individual judgments for inspection. Using 200 MICo-150K training tasks and 30 updates, the framework raises a GPT-5.4 composite score from 41.78 to 52.50 on a manually selected 897-task MICo-Bench subset; direct 27B rewards yield 51.84. Each reward is tested in one GRPO run, and offline human evaluation does not establish a statistically significant advantage over direct scoring. The study provides an initial implementation and evaluation of Visual Jev as a reference-bound reward for multi-subject image generation.