RFPO: Rectified Flow Policy Optimization for Embodied Control
This study addresses the degradation in control performance of flow policies under few-step execution caused by discretization errors, formally defining and resolving this "few-step discretization gap" for the first time. To this end, we propose the Rectified Flow Policy Optimization (RFPO) framework. RFPO corrects generation trajectories via reward-aware online reflow and integrates a Gaussian PPO controller for supervised optimization, enabling efficient and reliable robotic control under single-step Euler integration. Experimental results demonstrate that RFPO retains 98.5% of the full-step reward on the Go2 quadruped robot with single-step execution, reducing inference latency to 0.08 ms and achieving a 54.9× speedup. Furthermore, real-world deployment validates its stable control performance.