🤖 AI Summary
This study addresses the challenge of jointly optimizing reflection and generation in unified multimodal models, where conventional methods struggle to efficiently explore high-success repair paths. To this end, it proposes UMM-Reflection, a framework that end-to-end optimizes complete reflection trajectories within a single model via interleaved reinforcement learning. Specifically, it introduces trajectory-level advantage estimation to circumvent combinatorial explosion and integrates GRPO with flow matching, enabling cross-turn credit assignment and synchronous updates of diagnostic and corrective roles without external verifiers. This approach facilitates the co-evolution of image diagnosis and correction. Empirically, UMM-Reflection improves GenEval performance by 12.05 points on BAGEL and demonstrates substantial generalization across multiple unseen benchmarks, including WISE.
📝 Abstract
Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.