🤖 AI Summary
This study addresses the "task island" problem in Vision-Language-Action (VLA) models, wherein preceding tasks alter physical states and render subsequent instructions ineffective. To overcome this limitation, this work proposes a training-free inference-time intervention mechanism that requires neither parameter updates nor external modules. By backpropagating through a frozen decoder, the method performs state-guided writes within the action stream representations, enabling the VLA to autonomously recover its motion. The primary contribution lies in achieving cross-task state recovery relying solely on internal representation intervention. Evaluations on the LIBERO benchmark and real-world robot experiments demonstrate that task-switching success rates improve by 47%–65%, substantially expanding the task reachability of VLA models.
📝 Abstract
Vision-language-action (VLA) policies can execute many tasks from standard initial states, yet a new instruction may fail after another task has altered the robot's physical state. We study instruction switching, where a new task is issued during or after the execution of a different one. We observe that a target task that is reliably completed from its standard initial states can become inaccessible from states produced by a preceding task. We call such states task islands. We propose Causeway, a training-free inference-time intervention. Given the current state and a re-entry pose for the target task, Causeway back-propagates through the frozen decoding computation and applies a state-directed write within the action-stream representation. The VLA decodes the return motion itself, without parameter updates, a new action head, or external action generation. Across 71 cross-object pairs, three switch timings, and three VLA architectures on LIBERO-Goal, Causeway raises bare-switch success from 3-26% to 47-65% and increases the rate of reaching the handoff neighborhood by 42-72 percentage points across models. Additional experiments on LIBERO-Object and a real xArm platform show that the recovery extends beyond the main LIBERO-Goal setting, both in simulation and on a robot.