🤖 AI Summary
This study addresses the challenge in generative online reinforcement learning where sampling from unnormalized densities induces high variance in importance sampling and training instability. To this end, it proposes the Score-Calibrated Flow algorithm, which establishes score-calibrated optimality conditions and leverages a fixed-point equation of the velocity field to construct a stop-gradient objective for self-consistent optimization. This approach preserves the scalable architecture of conditional flow matching (CFM), enabling efficient training of generative policies without requiring target samples or backpropagation. Empirical evaluations on standard RL benchmarks demonstrate that the proposed method matches or surpasses state-of-the-art baselines in performance while significantly reducing training time and enhancing both sample efficiency and stability.
📝 Abstract
Diffusion and flow models provide expressive policy classes for online reinforcement learning (RL), enabling multimodal behaviors and improved performance. However, training these policies remains challenging: the critic specifies the desired policy as an unnormalized Boltzmann density but does not provide direct samples from it. Many existing methods rely on importance sampling to construct training signals, which can suffer from high variance, increasing computational cost and destabilizing training. We propose Score-Calibrated Flow (SCF), a simple and efficient algorithm for training generative models to sample from unnormalized densities without importance sampling or backpropagation through the sampling trajectory. We learn the desired flow by enforcing self-consistency, bypassing target posterior mean estimation. By jointly exploiting the prescribed target score and the structure of flow matching, we establish these self-consistency requirements as score-calibrated optimality conditions, first for the terminal density and then for the trainable velocity field. We prove that their unique solutions are, respectively, the target density and the ideal flow model that conditional flow matching (CFM) would recover if target samples were available. We formulate the velocity condition as a fixed-point equation and exploit its conditional-expectation structure to construct a stop-gradient objective for enforcing it. The resulting training procedure retains the scalable sample-interpolate-regress structure of CFM despite the absence of target samples, using endpoints generated by the current flow. For online RL, the critic gradient supplies the target score at the generated actions, yielding a direct approach to actor training. Experiments on RL benchmarks demonstrate that SCF matches or improves upon state-of-the-art generative-policy baselines, while substantially reducing training time.