🤖 AI Summary
This study addresses the challenge in online safe reinforcement learning where multimodal action distributions, induced by reward and safety constraints, often cause conventional Gaussian policies to collapse and destabilize Lagrangian optimization. To this end, we propose RAFALE, which formulates policy updates as a constrained generalized Schrödinger bridge problem, proving its equivalence to action-space entropy regularization and convergence to minimum-energy mappings as noise vanishes. By differentiating augmented objectives directly through flow-generated trajectories to circumvent score function estimation, RAFALE integrates an off-policy actor-critic architecture with flow matching and density-agnostic kinetic energy regularization for efficient safe updates. Evaluated across seven Safety-Gymnasium tasks, RAFALE maintains highly competitive rewards while strictly satisfying cost constraints, significantly outperforming mainstream baseline methods.
📝 Abstract
Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. Reward and safety can induce multimodal action distributions, challenging the prevailing primal-dual methods: Gaussian actors may collapse onto a single suboptimal mode, and optimization over the nonconvex Lagrangian landscape can be unstable. Diffusion and flow policies can represent such distributions, but recent work with a diffusion actor relies on estimating and matching the score of an augmented-Lagrangian target policy. Instead, we differentiate the augmented objective directly through the generation path of a flow policy, so no score needs to be estimated. Because a flow policy lacks a readily available action log-density for entropy regularization, we build on the density-free kinetic-energy regularizer of FLAC, a recent reward-only method, and propose Reparameterized Augmented-Lagrangian Flow Actor with Least Energy (RAFALE), an off-policy actor-critic method for safe RL. We formulate its update as a constrained one-ended generalized Schr\"odinger bridge and show that, for each source draw, this path-space problem is exactly an entropy-regularized problem in action space. At positive noise, its solution reweights the reward-only action distribution only where the estimated cost exceeds a threshold set by the Lagrange multiplier. As the noise vanishes, the optimal value converges to that of a least-energy map objective that the flow policy optimizes directly. Across seven Safety-Gymnasium tasks, RAFALE achieves competitive reward with mean final cost within budget on every task, whereas strong baselines trade one for the other; ablations support the necessity of both its augmented objective and its flow actor.