🤖 AI Summary
This study addresses the decision-making inefficiency caused by sampling in flow-based policy iteration. We propose an efficient offline reinforcement learning method based on mean velocity supervision, which reformulates adjoint signals into mean velocity targets to directly learn finite-interval transport, thereby enabling few-step action generation. Furthermore, this work presents the first integration of adjoint Q-optimization with the MeanFlow architecture, allowing the training of efficient Actor-Critic policies without backpropagating through sampled trajectories. On the HumanoidMaze benchmark, the proposed approach achieves high-performance decision-making with only two network evaluations, yielding results comparable to strong baselines.
📝 Abstract
Flow policies can model rich action distributions, but their iterative sampling limits decision speed. Adjoint matching uses the critic's action gradient to improve a flow policy without backpropagating through its sampling trajectory, yet its supervision is defined for instantaneous velocities. We propose QAMM, a method that turns the critic-derived adjoint signal into supervision for MeanFlow's average velocity. The resulting policy learns finite-interval transport directly and generates actions with few network evaluations. We derive the adjoint MeanFlow target, specify its gradient boundaries, and train it with an offline actor-critic. On ten HumanoidMaze tasks, QAMM produces effective two-call policies and achieves competitive performance against strong flow-policy baselines. These results show that adjoint-based Q optimization can be combined with average-velocity learning to obtain expressive offline policies with few-step action generation.