TraceFlow: Guiding Frozen Flow-Matching Robot Policies with Success and Failure Traces

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出TraceFlow方法,通过成功和失败轨迹修正冻结的流匹配动作专家,提高机器人任务成功率。
📝 Abstract
A vision-language-action (VLA) policy with a flow-matching action expert generates each action chunk (a short command sequence) by integrating a learned velocity field; once its weights are fixed, the success or failure of an earlier rollout cannot change the chunk generated now. Concurrent test-time methods give a frozen policy such an input from retrieved successes, a learned critic, a verifier, or a dynamics model, but none uses the robot's own failed rollouts as negative evidence with nothing but a terminal outcome bit. We introduce TraceFlow, a progress-aligned guidance field that turns the action densities of retrieved successful and failed rollouts into a bounded correction to a frozen flow-matching action expert, using one terminal outcome bit per rollout and no other label. Its TraceBank stores traces, time-ordered state-action records with a terminal label, starts from the target-task training traces, and later admits the deployed robot's own rollouts. On an ordered real-robot packing task the base completes 21 of 50 trials in order, TraceFlow 39, and one stacking round without any weight update 47, with wrong-sequence episodes falling from 20 to 0. In simulation the gain is selective: with per-suite selected settings, TraceFlow raises RoboMemArena Sequence from 78.92\% to 91.50\% task success and Transferring from 54.41\% to 62.00\% at stacking round 2, leaves the 26-task aggregate unchanged, lowers Counting and Occlusion by 1.12 and 1.42 points, and changes LIBERO-Plus (Long) by +1.27 points (p = 0.0733). Stacking gains are finite, every branch peaking before round ten, and the bank's success-to-failure ratio predicts no retrieval allocation.
Problem

Research questions and friction points this paper is trying to address.

robot policies
flow-matching
success and failure traces
frozen policy
action expert
Innovation

Methods, ideas, or system contributions that make the work stand out.

TraceFlow
flow-matching action expert
terminal outcome bit
TraceBank
J
Jiaxuan Zhang
The University of Hong Kong, Hong Kong SAR, China; Southern University of Science and Technology, Shenzhen, China; HKU InfoBodied AI Lab
R
Ruizhe Liu
The University of Hong Kong, Hong Kong SAR, China; HKU InfoBodied AI Lab
Y
Yu Zhang
The University of Hong Kong, Hong Kong SAR, China; HKU InfoBodied AI Lab
Yanchao Yang
Yanchao Yang
Assistant Professor, HKU; Stanford University; UCLA
Embodied AIComputer VisionMachine Learning