RecastVLA: From Past Interaction to Future Control with Adaptive Policy States

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that current observations in sequential robotic manipulation fail to effectively encode historical interaction information. To overcome this, it proposes an action-side test-time training flow matching strategy that pioneers the transformation of action generation features into persistent policy states. By leveraging a shared fast-weight mechanism to record historical context for control guidance, the approach enables online updates without expert labels and facilitates cross-depth interface sharing. The proposed method significantly improves success rates on benchmarks such as LIBERO and in real-world robotic tasks, notably achieving a 10.68 percentage point increase on RoboTwin.
📝 Abstract
Sequential manipulation requires a robot to track what has already happened, even when the current scene no longer reveals it. Policies with explicit history representations make past interactions available as context for current decisions. We ask how action generation itself can form a persistent state for subsequent control. Building on action-side test-time training, RecastVLA maintains an adaptive policy state within a flow-matching vision-language-action policy. The state is represented by shared fast weights and remains fixed throughout action generation. Depth-specific interfaces read the same state, while features across depths and flow evaluations jointly define one update for the next policy call. Subsequent action losses train the initialization, interfaces, and update rule by differentiating through earlier state transitions. At deployment, updates use the policy's own action-generation features without expert action labels. Across LIBERO, RoboTwin, RoboDojo, and twelve real-robot tasks, RecastVLA improves mean success over a matched policy trained without test-time training, including 10.68 percentage points on RoboTwin Clean-to-Clean. In controlled RoboTwin comparisons, retaining state improves success, and the shared design exceeds independently trained layer-local TTT by 2.58 points.
Problem

Research questions and friction points this paper is trying to address.

sequential manipulation
policy state
history representation
test-time training
vision-language-action
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Policy
Test-Time Training
Flow Matching
Fast Weights
Sequential Manipulation
🔎 Similar Papers
No similar papers found.
Wenbo Li
Wenbo Li
The Chinese University of Hong Kong
Computer VisionDeep Learning
J
Jun Yang
Yuanwu Technology, Shenzhen, China
Y
Yiteng Chen
School of Software Engineering, South China University of Technology, Guangzhou, China
W
Wei Zhang
School of Software Engineering, South China University of Technology, Guangzhou, China
Qingyao Wu
Qingyao Wu
School of Software Engineering, South China University of Technology
Computer VisionMachine Learning