rl fine-tuning for flows

Design and implement reinforcement-learning fine-tuning pipelines that modify flow-based generative models and their samplers—particularly models trained with flow-matching—to steer generation toward task-specific objectives (e.g., higher fidelity, greater diversity, or satisfaction of process-aware constraints). This work includes specifying reward functions, integrating RL agents with sampling procedures, choosing and tuning RL algorithms, and measuring how fine-tuning changes generative behavior.

rlfine-tuningforflows

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Feb 09, 2025
JF
Jiajun Fan
🏛️ University of Illinois Urbana-Champaign | Zhejiang University | Tsinghua University | University of Southern California

This paper addresses two key challenges in online reinforcement learning (RL) fine-tuning of continuous flow-based generative models: policy collapse and high computational cost of likelihood estimation. We propose a novel framework that requires neither reward gradients nor data filtering. Our method integrates online reward-weighted importance reweighting with Wasserstein-2 distance regularization into conditional flow matching (CFM), enabling tractable upper-bound computation and guaranteeing theoretical convergence. Crucially, our framework establishes, for the first time, an equivalence to KL-regularized RL algorithms—thereby jointly optimizing policy performance and generation diversity. Extensive experiments on target image generation, image compression, and text–image alignment demonstrate substantial improvements in reward scores while preserving high-fidelity distribution coverage.

Aligning models with arbitrary reward functionsFine-tuning continuous flow-based generative modelsPreventing policy collapse and maintaining diversity

Existing reward-based fine-tuning of generative models lacks theoretical foundations—particularly within flow matching and denoising diffusion frameworks—struggling to jointly ensure accuracy, sample diversity, and generalization to unseen human preferences. Method: This work formally recasts reward-driven generation as a stochastic optimal control (SOC) problem. We rigorously prove that memoryless noise scheduling is necessary to decouple noise from samples, thereby enabling tractable optimization. Building on this insight, we propose Adjoint Matching: a novel algorithm that transforms the SOC formulation into a supervised regression task via adjoint-state methods, unifying stochastic optimal control, adjoint calculus, flow matching, and diffusion modeling. Results: Experiments demonstrate that our approach significantly outperforms state-of-the-art methods in fidelity, perceptual realism, and generalization to unseen reward models, while preserving high sample diversity.

Denoising Diffusion ModelsReward MechanismStream Matching

This work addresses the challenge that offline reinforcement learning policies often struggle to adapt during online fine-tuning due to insufficient exploration. To bridge the gap between offline pretraining and online adaptation, the authors propose FINO, a novel approach that integrates flow matching generative models with action-space noise injection and an entropy-guided sampling mechanism. This combination explicitly enhances exploratory behavior within the pretrained policy, enabling efficient fine-tuning under limited online interaction budgets. By promoting structured exploration while preserving learned offline knowledge, FINO significantly improves sample efficiency. Empirical evaluations across diverse and complex tasks consistently demonstrate that FINO outperforms existing state-of-the-art methods.

explorationoffline-to-online reinforcement learningpolicy adaptation

This work addresses the collapse of generation diversity in reinforcement fine-tuning, where optimization dynamics often drive model outputs toward a single solution (i.e., a Dirac delta distribution) due to misalignment between the objective function and the optimization landscape. To mitigate this, we propose DRIFT, the first framework to systematically incorporate diversity incentives into reinforcement fine-tuning. DRIFT synergistically preserves both task alignment and output diversity during policy updates through reward-concentrated subset sampling, stochastic prompt augmentation, and potential-based reward shaping. Experimental results demonstrate that DRIFT achieves Pareto superiority: it improves generation diversity by 9.08%–43.46% while maintaining equivalent task alignment, or enhances task alignment by 59.65%–65.86% under comparable diversity levels.

Dirac delta distributiondiversity collapsegenerative models

GFlowNet Training by Policy Gradients

Aug 12, 2024
PN
Puhua Niu
🏛️ Texas A&M University

GFlowNets suffer from low training efficiency and unstable gradient estimation in combinatorial object generation due to strict flow conservation constraints. To address this, we propose the first policy-gradient-based GFlowNet training framework. Our method reformulates flow conservation as a policy optimization objective, enabling joint training of forward and backward policies without explicit flow matching. We provide theoretical convergence guarantees and introduce a coupled update mechanism to reduce gradient variance. Experiments across multiple synthetic and real-world datasets demonstrate that our approach significantly improves sample quality, training stability, and fidelity to the target distribution—particularly under sparse-reward settings, where it exhibits superior robustness.

Develops coupled strategy for joint forward and backward policy trainingImproves GFlowNet performance via robust gradient estimation techniquesProposes new GFlowNet training framework with policy-dependent rewards

Latest Papers

What's happening recently
View more

This work addresses the challenge of end-to-end optimization when combining flow matching with value-gradient-based reinforcement learning, a setting where sampling instability often undermines performance. Existing approaches typically compromise either representational capacity or the iterative generative nature of the process. To overcome this, we propose VINE, a stable sampling method tailored for reinforcement learning that constructs differentiable trajectories by reconstructing interpolated states at each denoising step. Our analysis reveals that the observed instability stems from the original sampling strategy rather than the iterative generation mechanism itself. VINE is the first method to enable end-to-end value-gradient optimization while preserving high expressivity, achieving substantial improvements over state-of-the-art baselines on both the OGBench offline reinforcement learning benchmark and real-world robotic manipulation tasks.

flow-matchinggenerative control policiesiterative generation

This work addresses the limitations of large language models in generating high-quality BPMN process models, which are constrained by supervised fine-tuning data and the absence of well-defined multidimensional reward functions. The authors propose a reinforcement learning–based optimization approach that systematically explores a reward function encompassing 38 syntactic, pragmatic, and semantic metrics. They train Llama-3.1-8B and Qwen2.5-14B models across 48 configurations and find that uniformly weighted rewards outperform targeted weighting schemes, with significant interaction effects observed between reward composition and model architecture. Leveraging Group Relative Policy Optimization and an automated evaluation framework, the method substantially improves pragmatic and syntactic quality while preserving semantic fidelity and reducing output variability by over sixfold. All code is publicly released.

LLMmulti-dimensional qualityprocess model generation

This work addresses the challenges of exploration and denoising stability in reinforcement learning post-training with flow matching models by introducing Precise Sampler—a stochastic sampling strategy consistent with stochastic differential equations (SDEs). The method approximates the clean latent posterior mean via a frozen estimate, thereby eliminating redundant noise inherent in standard discretization schemes and effectively balancing exploration and stability within a limited number of sampling steps. Integrating SDE scheduling, flow matching theory, and reinforcement learning, the proposed approach achieves state-of-the-art performance on benchmarks such as PickScore and HPSv2.1, while reducing training time by 13.1%–53.2%, significantly enhancing both the efficiency and convergence stability of reward optimization.

denoising trajectoryflow-matching modelsreinforcement learning

Hot Scholars

YM

Yichuan Ma

Fudan University
LLMSynthetic Data
SZ

Shi-Zhe Chen

Tencent
Computer VisionNatural Language Processing