CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of manual window specification, reward saturation, and sample inefficiency in reinforcement learning for diffusion models by proposing CAST. This method introduces the first denoising-trajectory-based adaptive SDE sampling window selection. Furthermore, it incorporates causal scene graphs to decompose prompts into verifiable atomic units, thereby eliminating reward saturation. To enable fine-grained fine-tuning, CAST projects the advantage function into pixel space and integrates a teacher-forcing attention mechanism. Evaluated on the GenEval 2 benchmark, CAST achieves an improvement margin on challenging prompts that is 3.07 times greater than that of Flow-GRPO, while significantly enhancing overall generation quality.
📝 Abstract
Online reinforcement learning has been extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where exploration noise is injected. We instead determine it from each model's denoising trajectory. (2) Reward saturation. Current methods rely on scoring models trained on human annotations; we find that such scores are extremely high and nearly indistinguishable on the latest SOTA open-source DMs, making advantage estimation largely ineffective. (3) Sample inefficiency. A single scalar reward collapses different failure modes into almost identical scores, leaving minimal gradient guidance for targeted improvement. To address these issues, we propose CAST (Causal Advantage-Structured Training), an RL fine-tuning method for pretrained DMs, which (1) identifies the denoising step at which each model fixes the objects and their spatial arrangement in the image and uses that timing to set the SDE window, (2) decomposes each prompt via Causal Scene Graphs (CSG) into verifiable-atoms, i.e., minimal semantic units such as an object, count, attribute, or spatial relation that can each be checked independently, and rewards each atom separately, and (3) projects the signed atom-level advantages into pixel space through teacher-forced attention and uses them to spatially weight the SDE policy objective. We fine-tune two of the strongest open-source DMs, FLUX.2-dev and Qwen-Image-2512, with CAST, and evaluate them on GenEval 2, a compositional benchmark, and on Qwen-Image-Bench for overall quality. Within almost the same training budget, CAST's improvement over the base model on the most challenging GenEval 2 prompts is up to 3.07x that of Flow-GRPO, while overall generation quality also improves.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Models
Reinforcement Learning
Reward Saturation
Sample Inefficiency
Compositional Image Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Models
Reinforcement Learning
Causal Scene Graphs
Compositional Rewards
Flow Matching
💼 Related Jobs
No related jobs found.
S
Shu Yu
Shanghai Artificial Intelligence Laboratory, Shanghai, China; Shanghai Innovation Institute, Shanghai, China; Fudan University, Shanghai, China
Chaochao Lu
Chaochao Lu
Shanghai AI Laboratory
Causal AI