🤖 AI Summary
This study addresses the challenges of local errors caused by distribution shift and reward scarcity in real-world scenarios during the deployment of flow-matching robotic policies. To this end, it proposes a test-time guidance framework based on QGF sampling. The method leverages human intervention data to train a preference model and innovatively introduces gradient upper bounds alongside zero-gradient losses to regularize the direction of preference gradients, thereby optimizing pretrained policies without updating the base model or requiring environmental rewards. Evaluated across four real-world precision insertion tasks, the proposed approach achieves an average success rate of 90.5%, representing a substantial improvement over the 69% attained by frozen policies. These results validate the effectiveness of utilizing intervention data as localized supervision for enhancing policy performance at test time.
📝 Abstract
Flow-matching policies can represent complex robot behaviors but remain susceptible to local errors under distribution shift at deployment. Many reinforcement learning approaches to policy improvement require reward signals that are difficult to specify or obtain in real-world manipulation. We present PreferenceFlow, a framework for improving a pretrained flow policy at test time without environment rewards or updates to the base policy. Human intervention chunks are paired with robot chunks generated from the same initial conditioning state to train a preference model. During infer- ence, we adopt the QGF sampling update, replacing its value gradient with the preference gradient evaluated at an estimated clean action. A gradient-cap loss penalizes excessive gradients on intervention pairs, while a zero-gradient loss discourages guidance near actions from expert demonstrations. On four real-world precision insertion tasks with a Franka robot, Pref- erenceFlow achieves a mean success rate of 90.5%, compared with 69% for the frozen policy. Ablations support the roles of gradient regularization and correctly ordered preference labels in the evaluated settings. These results demonstrate the utility of human interventions as local preference supervision for guiding frozen generative robot policies.