🤖 AI Summary
This study addresses the suboptimal policy updates in Flow-GRPO for image generation, where uniform scalar advantages overlook spatial structures. To overcome this limitation, we propose a spatial gradient-guided credit assignment framework tailored for diffusion Transformers (DiTs). Methodologically, fine-grained alignment is achieved through token-level reconstruction and reward-gradient-based continuous spatial credit maps, while Median Absolute Deviation (MAD) robust normalization is introduced to suppress noise and highlight critical regions. Experimental results demonstrate that, when integrated with importance sampling and temperature scaling, our approach attains state-of-the-art alignment quality on the GenEval benchmark. Furthermore, it achieves convergence speeds comparable to DiffusionNFT and significantly outperforms existing Flow-GRPO methods, validating the effectiveness of spatially aware credit assignment in reinforcement learning for diffusion models.
📝 Abstract
Reinforcement Learning (RL) has proven effective in aligning flow-based generative models with human preferences. Recently, Flow-GRPO has emerged as an efficient critic-free paradigm by calculating advantages over sampled candidate trajectories. However, standard Flow-GRPO applies a uniform scalar advantage across both temporal denoising steps and spatial latent dimensions, without explicitly accounting for the spatial structure of generated images, which may lead to sub-optimal policy updates. To address this, we propose a novel gradient-guided spatial credit assignment framework tailored for Diffusion Transformers (DiTs). We first reformulate the transition-level log-likelihood in Flow-GRPO into a token-wise representation natively aligned with DiT patch architectures, constructing spatially fine-grained importance sampling ratios. To allocate localized credit without rigid, boundary-sensitive segmentation heuristics, we introduce a continuous spatial credit map derived from reward gradients. Crucially, we employ an outlier-robust normalization scheme based on Median Absolute Deviation (MAD) coupled with temperature scaling, effectively eliminating gradient noise while highlighting functional prompt-aligned regions. Extensive evaluations on GenEval show that our approach delivers SOTA alignment quality, achieving a convergence rate comparable to top-tier methods like DiffusionNFT while substantially improving upon Flow-GRPO-based methods in alignment performance.