RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy

πŸ“… 2026-07-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the error accumulation in flow-matching vision-language-action (VLA) policies under offline reinforcement learning, which arises from distributional shift and is exacerbated by existing methods that only utilize failed trajectories at the trajectory levelβ€”leading to inefficient learning and limited error correction. To overcome this, we propose RedFlow, a framework that enables fine-grained, action-level exploitation of failure experiences for the first time. RedFlow employs a context-aware corrective matching mechanism to identify erroneous actions and retrieve similar successful alternatives, combined with an adaptive redirection objective that transforms mixed-quality data into dense supervision signals. Evaluated on the LIBERO benchmark and three real-world manipulation tasks, RedFlow boosts success rates from 56.7% to 74.7%, matching the performance of strong on-policy methods such as PPO and GRPO while using nearly an order of magnitude fewer training samples.
πŸ“ Abstract
Flow-matching Vision-Language-Action (VLA) policies have shown strong potential for robotic manipulation but often suffer from compounding errors caused by distribution shifts during deployment. While offline reinforcement learning (RL) provides a practical way to improve deployed policies using rollout data, existing methods either ignore failure data or exploit it only at the trajectory level, resulting in low learning efficiency and persistent errors. We propose **RedFlow**, a fine-grained offline RL framework that redirects failure experiences into action-level corrective supervision for flow-matching VLA policies. RedFlow consists of two key components: (1) a **Context-Aware Corrective Matching** mechanism that identifies failure-inducing actions and retrieves successful alternatives from similar contexts as corrective targets, and (2) an **Adaptive Redirection Objective** that jointly reinforces successful actions, suppresses undesirable ones, and redirects recoverable failures toward corrective targets. By converting both successful and failed experiences into dense supervision, RedFlow enables robust recovery learning from mixed-quality data. Experiments on the LIBERO benchmark and three real-world manipulation tasks show that RedFlow consistently outperforms state-of-the-art offline RL baselines, improving the real-world success rate from 56.7% to 74.7%. It also matches strong on-policy methods (PPO, GRPO, and DDPO) while requiring roughly an order of magnitude fewer training samples.
Problem

Research questions and friction points this paper is trying to address.

flow-matching
Vision-Language-Action
offline reinforcement learning
distribution shift
failure correction
Innovation

Methods, ideas, or system contributions that make the work stand out.

RedFlow
flow-matching
offline reinforcement learning
action-level correction
VLA policy
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
Z
Zhengyang Yan
The Hong Kong University of Science and Technology
Junhao Li
Junhao Li
Assistant Project Scientist, Cognitive Science, University of California, San Diego
Non-coding RNAsDNA methylationEpigeneticsBioinformatics
F
Fangqi Zhu
The Hong Kong University of Science and Technology
Z
Zijun Wang
The Hong Kong University of Science and Technology
Q
Quanxin Shou
The Hong Kong University of Science and Technology
Y
Yikun Miao
The Hong Kong University of Science and Technology
X
Xiaoyi Pang
The Hong Kong University of Science and Technology
Zicong Hong
Zicong Hong
Department of Computer Science and Engineering, Hong Kong University of Science and Technology
BlockchainML SystemEdge/Cloud Computing
Song Guo
Song Guo
Chair Professor of CSE, HKUST
Large Language ModelEdge AIMachine Learning Systems