π€ AI Summary
This study addresses the performance limitations of existing Vision-Language-Action (VLA) models on multi-stage tasks caused by the absence of long-term memory. To this end, it proposes a State Evolution Memory mechanism that explicitly encodes historical state transitions to track task progress. The core methodology comprises Conditional Delta Tokenization for precisely characterizing inter-frame dynamics, a shared VLM-based task-adaptive routing mechanism that intelligently distinguishes between routine and multi-stage tasks, and a joint training strategy enabling end-to-end optimization. Experimental results demonstrate that the proposed approach achieves success rates of 80.7%, 82.0%, and 83.8% on RMBench, RoboMME, and real-world robotic tasks, respectively, significantly outperforming existing state-of-the-art methods.
π Abstract
Most vision-language-action (VLA) models rely on current observations and lose task-relevant evidence once it leaves view, limiting performance on long-horizon, memory-dependent tasks. Existing efforts incorporate compressed historical features or sparse visual keyframes. However, isolated snapshots can leave the policy uncertain about what changed during past interactions and which action should follow. To overcome this limitation, we propose EvoMem-VLA, which constructs state-evolution memory by explicitly encoding and retaining observed changes between historical states. These change representations preserve evidence of interaction outcomes, allowing the policy to track task progress beyond isolated snapshots. Specifically, we introduce conditional delta tokenization to encode ordered frame pairs into directional, source-conditioned delta tokens, each associated with its corresponding state evidence. A shared VLM backbone supports task-adaptive routing: normal long-horizon tasks follow a direct action route, whereas multi-stage tasks use a subtask route that generates an executable subtask as an additional input for action generation. With a single jointly trained policy for each simulation benchmark, EvoMem-VLA achieves success rates of 80.7\% on RMBench, 82.0\% on RoboMME and 83.8\% across four real-world tasks spanning two robot embodiments. These results represent substantial improvements over the previous state of the art in all three evaluation settings.