🤖 AI Summary
This study addresses the inefficiency of visuomotor policy generation arising from multi-step integration and its neglect of geometric constraints inherent to robot action manifolds. To this end, we propose Riemannian MeanFlow policies. This method introduces Riemannian conditional flow matching coupled with a flow-map consistency objective, leveraging Riemannian anchors to learn conditional flow maps directly on the action manifold. This formulation guarantees finite-time transport and enables single-step sampling. Evaluated on the LASA and Robomimic benchmarks as well as real-world robotic tasks, the proposed approach achieves competitive performance at significantly reduced sampling costs. Consequently, this work establishes a new paradigm for motion planning that simultaneously ensures geometric consistency and highly efficient policy generation.
📝 Abstract
Visuomotor policies learn a direct map from raw sensory observations to robot action sequences. Policies based on Diffusion and Flow Matching capture the multimodal distribution over action sequences in an end-to-end manner. This expressivity comes at the cost of multi-step numerical integration of the learned vector field for action generation, which can be expensive and time-consuming, impeding fast control rates required in robotics applications. Furthermore, robot action sequences are usually defined on a smooth, differentiable manifold, requiring that the learned policy respects the intrinsic geometry of the robot's action space. Here, we present Riemannian MeanFlow Policy (RMFP), which learns the conditioned flow map of the probability path on the robot action manifold. Our formulation employs a flow map consistency objective grounded in the data by a Riemannian Conditional Flow Matching anchor. The flow map consistency condition is stable to train and constrains the learned model to finite-time transport, which yields on-manifold action sequence generation with as few as one network function evaluation. We present results on the spherical LASA and Push-T benchmarks, on the Tool Hang and Transport tasks of the Robomimic suite, and on the Franka Kitchen task with manifold-constrained action generation, and demonstrate that RMFP attains performance competitive with prior work at a lower sampling cost. We also employ RMFP on a real-world robotic manipulation task to demonstrate fast action generation under imperfect sensor measurements in the physical world.