Fewer Steps, Better Actions: Rethinking Flow-Matching Inference for VLA Policies

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究提出Coda方法,通过减少步骤并引入端点校正提高VLA策略的性能和效率,解决了增加集成步骤不必然提升闭环成功的问题。
📝 Abstract
Vision-language-action (VLA) policies based on flow matching generate action chunks through repeated evaluations of an action expert. Increasing the number of integration steps raises inference cost, but does not necessarily improve closed-loop success. We propose Coda, which reallocates part of this integration budget to a single learned endpoint correction. A frozen policy first completes a few-step noise-to-action trajectory; a lightweight Transformer then predicts a demonstration-supervised residual using the candidate action, source noise, and shared observation-prefix cache. Only the corrector is trained. On 50 RoboTwin Easy tasks, five-step Coda improves success from 71.64% to 74.68% over the matched five-step baseline, while reducing forward latency by 30.2% relative to the default ten-step policy. A two-step configuration achieves 71.88% success with a 2.12$\times$ speedup. An independent 13-task control shows a 5.69-percentage-point gain at nearly equal latency, supporting correction as an effective alternative to additional integration. The same design also improves frozen official SmolVLA, raising two-step success from 60.8% to 69.4%. These results show that endpoint correction improves the quality-latency trade-off of frozen flow-matching policies.
Problem

Research questions and friction points this paper is trying to address.

Vision-language-action
Flow matching
Inference cost
Closed-loop success
Innovation

Methods, ideas, or system contributions that make the work stand out.

Coda
Endpoint Correction
Flow Matching
Transformer
Latency Reduction
Zhipeng Tang
Zhipeng Tang
UMass Amherst
X
Xinda Chen
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
W
Weining Rao
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
X
Xiao Li
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
W
Wenting Tan
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
Y
Yuning Wang
Zhongke Haichuan Intelligent
X
Xiao Shi
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
X
Xiaofang Zhao
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China