RFPO: Rectified Flow Policy Optimization for Embodied Control

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation in control performance of flow policies under few-step execution caused by discretization errors, formally defining and resolving this "few-step discretization gap" for the first time. To this end, we propose the Rectified Flow Policy Optimization (RFPO) framework. RFPO corrects generation trajectories via reward-aware online reflow and integrates a Gaussian PPO controller for supervised optimization, enabling efficient and reliable robotic control under single-step Euler integration. Experimental results demonstrate that RFPO retains 98.5% of the full-step reward on the Go2 quadruped robot with single-step execution, reducing inference latency to 0.08 ms and achieving a 54.9× speedup. Furthermore, real-world deployment validates its stable control performance.
📝 Abstract
Flow-based policies provide an expressive framework for continuous robot control, but their iterative ODE integration incurs substantial inference cost. Naively reducing the integration budget can severely degrade control, since policies optimized under full-step execution are not explicitly constrained to remain reliable under coarse numerical integration. We refer to this mismatch as the few-step discretization gap. To address this problem, we introduce RFPO, a flow-policy optimization framework for reliable few-step execution. Reward-aware online Reflow rectifies student-induced transport paths during on-policy learning, making the resulting policy more robust to coarse integration. A frozen Gaussian PPO controller supplies complementary action-space supervision at full and intermediate integration budgets, while the deployed policy remains a single flow student executed with one Euler step. Across Unitree Go2, Boston Dynamics Spot, Unitree H1, and Unitree G1, RFPO consistently preserves full-step control performance under one-step execution, with one-step returns remaining within 2.4% of their corresponding 64-step values across both zero and random initialization. On Unitree Go2, one-step execution retains 98.5% of the 64-step reward while reducing onboard mean inference latency from 4.39 ms to 0.08 ms, yielding a 54.9x speedup. Real-robot experiments further validate stable one-step locomotion. Code: https://github.com/AIGeeksGroup/RFPO. Website: https://aigeeksgroup.github.io/RFPO.
Problem

Research questions and friction points this paper is trying to address.

flow-based policy
embodied control
inference cost
few-step discretization gap
continuous robot control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rectified Flow Policy Optimization
Few-step Discretization Gap
Reward-aware Reflow
Embodied Control
Inference Acceleration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Ting Huang
School of Computer Science, Peking University
L
Lisiyu Pan
School of Computer Science, Peking University
H
Haoyu Wang
School of Computer Science, Peking University
Zeyu Zhang
Zeyu Zhang
Gaoling School of Artificial Intelligence, Renmin University of China
LLM-based AgentResponsible RecSysCausal Learning
S
Siyuan Qian
School of Computer Science, Peking University
Y
Yanjun Li
School of Computer Science, Peking University
Y
Yandong Guo
AI2 Robotics
Boxin Shi
Boxin Shi
Peking University
Computer VisionComputational Photography
Hao Tang
Hao Tang
Peking University
computer vision