Score
Design and implement reinforcement-learning fine-tuning pipelines that modify flow-based generative models and their samplers—particularly models trained with flow-matching—to steer generation toward task-specific objectives (e.g., higher fidelity, greater diversity, or satisfaction of process-aware constraints). This work includes specifying reward functions, integrating RL agents with sampling procedures, choosing and tuning RL algorithms, and measuring how fine-tuning changes generative behavior.
Diffusion models achieve high generation quality but face significant challenges in aligning outputs with human preferences and safety constraints. This paper systematically surveys reinforcement learning– and reward modeling–based alignment methods for diffusion models, covering techniques such as human/AI hybrid feedback fine-tuning, direct preference optimization (DPO), and differentiable reward modeling. We propose the first three-dimensional taxonomy—“feedback type–optimization mechanism–safety efficacy”—to unify and characterize existing approaches. Five promising research directions are identified: multi-objective alignment, efficient active feedback acquisition, adversarial-robust safety alignment, online continual alignment, and interpretable image reward modeling. Through structured comparative analysis and comprehensive literature review, we delineate performance boundaries of current methods, providing both theoretical foundations and actionable research pathways toward building safe, trustworthy, and value-aligned generative AI systems.
This paper addresses two key challenges in online reinforcement learning (RL) fine-tuning of continuous flow-based generative models: policy collapse and high computational cost of likelihood estimation. We propose a novel framework that requires neither reward gradients nor data filtering. Our method integrates online reward-weighted importance reweighting with Wasserstein-2 distance regularization into conditional flow matching (CFM), enabling tractable upper-bound computation and guaranteeing theoretical convergence. Crucially, our framework establishes, for the first time, an equivalence to KL-regularized RL algorithms—thereby jointly optimizing policy performance and generation diversity. Extensive experiments on target image generation, image compression, and text–image alignment demonstrate substantial improvements in reward scores while preserving high-fidelity distribution coverage.
Existing reward-based fine-tuning of generative models lacks theoretical foundations—particularly within flow matching and denoising diffusion frameworks—struggling to jointly ensure accuracy, sample diversity, and generalization to unseen human preferences. Method: This work formally recasts reward-driven generation as a stochastic optimal control (SOC) problem. We rigorously prove that memoryless noise scheduling is necessary to decouple noise from samples, thereby enabling tractable optimization. Building on this insight, we propose Adjoint Matching: a novel algorithm that transforms the SOC formulation into a supervised regression task via adjoint-state methods, unifying stochastic optimal control, adjoint calculus, flow matching, and diffusion modeling. Results: Experiments demonstrate that our approach significantly outperforms state-of-the-art methods in fidelity, perceptual realism, and generalization to unseen reward models, while preserving high sample diversity.
This work addresses the challenge that offline reinforcement learning policies often struggle to adapt during online fine-tuning due to insufficient exploration. To bridge the gap between offline pretraining and online adaptation, the authors propose FINO, a novel approach that integrates flow matching generative models with action-space noise injection and an entropy-guided sampling mechanism. This combination explicitly enhances exploratory behavior within the pretrained policy, enabling efficient fine-tuning under limited online interaction budgets. By promoting structured exploration while preserving learned offline knowledge, FINO significantly improves sample efficiency. Empirical evaluations across diverse and complex tasks consistently demonstrate that FINO outperforms existing state-of-the-art methods.
This work addresses the collapse of generation diversity in reinforcement fine-tuning, where optimization dynamics often drive model outputs toward a single solution (i.e., a Dirac delta distribution) due to misalignment between the objective function and the optimization landscape. To mitigate this, we propose DRIFT, the first framework to systematically incorporate diversity incentives into reinforcement fine-tuning. DRIFT synergistically preserves both task alignment and output diversity during policy updates through reward-concentrated subset sampling, stochastic prompt augmentation, and potential-based reward shaping. Experimental results demonstrate that DRIFT achieves Pareto superiority: it improves generation diversity by 9.08%–43.46% while maintaining equivalent task alignment, or enhances task alignment by 59.65%–65.86% under comparable diversity levels.
GFlowNets suffer from low training efficiency and unstable gradient estimation in combinatorial object generation due to strict flow conservation constraints. To address this, we propose the first policy-gradient-based GFlowNet training framework. Our method reformulates flow conservation as a policy optimization objective, enabling joint training of forward and backward policies without explicit flow matching. We provide theoretical convergence guarantees and introduce a coupled update mechanism to reduce gradient variance. Experiments across multiple synthetic and real-world datasets demonstrate that our approach significantly improves sample quality, training stability, and fidelity to the target distribution—particularly under sparse-reward settings, where it exhibits superior robustness.
This work addresses the challenge of end-to-end optimization when combining flow matching with value-gradient-based reinforcement learning, a setting where sampling instability often undermines performance. Existing approaches typically compromise either representational capacity or the iterative generative nature of the process. To overcome this, we propose VINE, a stable sampling method tailored for reinforcement learning that constructs differentiable trajectories by reconstructing interpolated states at each denoising step. Our analysis reveals that the observed instability stems from the original sampling strategy rather than the iterative generation mechanism itself. VINE is the first method to enable end-to-end value-gradient optimization while preserving high expressivity, achieving substantial improvements over state-of-the-art baselines on both the OGBench offline reinforcement learning benchmark and real-world robotic manipulation tasks.
This work addresses the limitations of large language models in generating high-quality BPMN process models, which are constrained by supervised fine-tuning data and the absence of well-defined multidimensional reward functions. The authors propose a reinforcement learning–based optimization approach that systematically explores a reward function encompassing 38 syntactic, pragmatic, and semantic metrics. They train Llama-3.1-8B and Qwen2.5-14B models across 48 configurations and find that uniformly weighted rewards outperform targeted weighting schemes, with significant interaction effects observed between reward composition and model architecture. Leveraging Group Relative Policy Optimization and an automated evaluation framework, the method substantially improves pragmatic and syntactic quality while preserving semantic fidelity and reducing output variability by over sixfold. All code is publicly released.
This work addresses the challenges of exploration and denoising stability in reinforcement learning post-training with flow matching models by introducing Precise Sampler—a stochastic sampling strategy consistent with stochastic differential equations (SDEs). The method approximates the clean latent posterior mean via a frozen estimate, thereby eliminating redundant noise inherent in standard discretization schemes and effectively balancing exploration and stability within a limited number of sampling steps. Integrating SDE scheduling, flow matching theory, and reinforcement learning, the proposed approach achieves state-of-the-art performance on benchmarks such as PickScore and HPSv2.1, while reducing training time by 13.1%–53.2%, significantly enhancing both the efficiency and convergence stability of reward optimization.