🤖 AI Summary
This study addresses the high inference latency of multi-step flow matching in robot foundation models and the performance collapse incurred by directly applying MeanFlow. To overcome these challenges, we propose Kinematic MeanFlow (K-MF), a framework that introduces an intermediate-point decoupling mechanism grounded in kinematic identities. By isolating temporal derivative terms to independently model early- and late-stage denoising dynamics, K-MF effectively suppresses error amplification and resolves stability failures caused by local acceleration surges and magnitude distribution diffusion in velocity fields. Integrated with single-step action generation, K-MF reduces action head latency by 67.5%–74.4% and end-to-end latency by 30.3%–54.9% on the GR00T-N1.6 action head, demonstrating substantial efficiency gains without compromising stability.
📝 Abstract
In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a promising framework for this goal, yet its direct application leads to performance collapse. We discover that this stems from two distinctive dynamics exhibited in the RFM velocity field: (1) the ``local acceleration" exhibits stability early on, but surges sharply towards the end of the denoising process, and (2) the spread of its magnitudes across samples widens as denoising progresses. To address these issues, we introduce Kinematic MeanFlow (K-MF), a novel one-step action policy tailored for RFMs. Specifically, grounded in a kinematic identity, K-MF decouples the time derivative term in the MeanFlow formulation into two sub-interval terms separated by an intermediate point. This decoupled formulation enables the two terms to capture early-stage and late-stage denoising dynamics, respectively, while mitigating the error amplification across the process. As a result, our K-MF empowers RFMs to achieve one-step action generation in both training from scratch and fine-tuning paradigms across diverse tasks, while outperforming multi-step flow matching in most settings. In terms of inference efficiency, K-MF reduces action-head latency of GR00T-N1.6 by 67.5%~74.4% across L40 and Jetson Orin in eager and compiled modes, yielding end-to-end latency reductions of 30.3%~54.9%. Code will be available at https://github.com/IntelChina-AI/K-MF.