ExploreNet: Learning Where to Explore in Diffusion GRPO

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficient exploration caused by isotropic noise in reinforcement learning for diffusion models by proposing ExploreNet. Built upon Stable Diffusion 3.5 and Group Relative Policy Optimization (GRPO), the method introduces an adaptive noise prediction network that dynamically generates optimal noise scales for each latent variable. This work is the first to demonstrate the learnability of exploration distributions, revealing that their shape characteristics are more beneficial than mere magnitude adjustments, while incurring no additional computational overhead during inference. Experimental results show that ExploreNet achieves a 14% improvement on the GenEval2 benchmark and attains a 67.2% human preference win rate, significantly outperforming existing baseline methods.
📝 Abstract
Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This noise decides which rollouts the model learns from, yet it perturbs every channel and spatial position of the latent equally. In this paper, we instead show that latent elements differ in how much they change the generated image, so exploration should adapt to these differences. We introduce EXPLORENET to learn an adaptive exploration distribution. EXPLORENET is a policy that predicts a noise scale for every latent element from the current latent, the denoising step, and the prompt, before any reward is observed; it is trained on the reward spread of each rollout group and discarded after training, leaving inference unchanged. On Stable Diffusion 3.5 Medium, EXPLORENET improves held-out GenEval2 by 14% over Flow-GRPO, transfers to two independent compositional benchmarks and five preference and image-quality models, and reaches a 67.2% human preference win-rate. Overall, across our group-relative diffusion RL experiments, we find that exploration is learnable, the shape of the exploration distribution outweighs its magnitude, and rollout quality is more effective than rollout quantity.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Models
Reinforcement Learning
Exploration Strategy
Group-relative Policy Optimization
Latent Space Perturbation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Exploration
Diffusion GRPO
ExploreNet
Group-relative RL
Noise Scale Prediction
🔎 Similar Papers
2023-09-20IEEE transactions on circuits and systems for video technology (Print)Citations: 0