MEND: RL For Flow Models via Proximal Velocity Matching

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of KL penalties and the displacement cost mismatch during post-training reward alignment for flow models. To overcome these limitations, this work proposes a reinforcement learning alignment method based on proximal velocity matching. The approach discards the conventional KL term, freezes both the reference model and advantage weights, and introduces a quadratic displacement cost mechanism that accepts only gradient updates whose rewards exceed the associated displacement costs, thereby enabling efficient and generalizable flow model optimization. Experimental results demonstrate that the proposed method significantly outperforms baselines such as Flow-GRPO within very few update steps, achieving a PickScore of 24.03 and substantially surpassing ReFL and DiffusionNFT.
📝 Abstract
Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets, with no KL term, frozen reference model, or advantage weights. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol, it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A 300-update three-reward run also surpasses the five-reward DiffusionNFT model on all three rewards it trains on. MEND is general and easy to adopt: it applies to any flow backbone with a differentiable reward.
Problem

Research questions and friction points this paper is trying to address.

flow models
reward post-training
reinforcement learning
sample efficiency
proximal velocity matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proximal Velocity Matching
Flow Models
Reinforcement Learning
Reward Post-training
Sample Efficiency
🔎 Similar Papers
No similar papers found.