EvoMem-VLA: State-Evolution Memory for Long-Horizon Robot Manipulation

πŸ“… 2026-10-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the performance limitations of existing Vision-Language-Action (VLA) models on multi-stage tasks caused by the absence of long-term memory. To this end, it proposes a State Evolution Memory mechanism that explicitly encodes historical state transitions to track task progress. The core methodology comprises Conditional Delta Tokenization for precisely characterizing inter-frame dynamics, a shared VLM-based task-adaptive routing mechanism that intelligently distinguishes between routine and multi-stage tasks, and a joint training strategy enabling end-to-end optimization. Experimental results demonstrate that the proposed approach achieves success rates of 80.7%, 82.0%, and 83.8% on RMBench, RoboMME, and real-world robotic tasks, respectively, significantly outperforming existing state-of-the-art methods.
πŸ“ Abstract
Most vision-language-action (VLA) models rely on current observations and lose task-relevant evidence once it leaves view, limiting performance on long-horizon, memory-dependent tasks. Existing efforts incorporate compressed historical features or sparse visual keyframes. However, isolated snapshots can leave the policy uncertain about what changed during past interactions and which action should follow. To overcome this limitation, we propose EvoMem-VLA, which constructs state-evolution memory by explicitly encoding and retaining observed changes between historical states. These change representations preserve evidence of interaction outcomes, allowing the policy to track task progress beyond isolated snapshots. Specifically, we introduce conditional delta tokenization to encode ordered frame pairs into directional, source-conditioned delta tokens, each associated with its corresponding state evidence. A shared VLM backbone supports task-adaptive routing: normal long-horizon tasks follow a direct action route, whereas multi-stage tasks use a subtask route that generates an executable subtask as an additional input for action generation. With a single jointly trained policy for each simulation benchmark, EvoMem-VLA achieves success rates of 80.7\% on RMBench, 82.0\% on RoboMME and 83.8\% across four real-world tasks spanning two robot embodiments. These results represent substantial improvements over the previous state of the art in all three evaluation settings.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action (VLA)
Long-Horizon Manipulation
State-Evolution Memory
Robot Manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

State-Evolution Memory
Conditional Delta Tokenization
Task-Adaptive Routing
Vision-Language-Action Model
Long-Horizon Manipulation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Y
Yuheng Na
Xi’an Jiaotong University
Zhide Zhong
Zhide Zhong
Beijing Institute of Technology
Robotics
Junjie He
Junjie He
Guizhou University
MRIDeep LearningCT
J
Junfeng Li
The Hong Kong University of Science and Technology (Guangzhou)
Haodong Yan
Haodong Yan
PhD student of INTR, HKUST (GZ)
Human reconstructionmotion prediction
Jiaan Wang
Jiaan Wang
WeChat AI, Tencent
Natural Language ProcessingMachine TranslationInformation Systems
J
Jiaguan Zhu
The Hong Kong University of Science and Technology (Guangzhou)
Y
Yangyang Zheng
The Hong Kong University of Science and Technology (Guangzhou)
T
Tianyu Huang
The Chinese University of Hong Kong
Haoang Li
Haoang Li
Assistant Professor, Hong Kong University of Science and Technology (Guangzhou)
Robotics3D Computer Vision