🤖 AI Summary
This study addresses the challenge of evaluating long-term memory in Vision-Language-Action (VLA) models, which typically observe only recent frames. We propose a novel memory-dependent task design paradigm based on information gaps, where critical cues are concealed to compel models to retrieve historical memory for successful manipulation. Accordingly, we construct a large-scale benchmark dataset comprising 90 tasks and 22,500 trajectories, stored in standard RLDS and LeRobotDataset v3 formats, with oracle trajectories provided to quantify memory efficacy. Fine-tuning experiments using the π0.5 baseline model establish a success rate baseline of 0.211±0.044 without explicit memory modules. This work provides a standardized evaluation platform for advancing long-term memory research in embodied intelligence.
📝 Abstract
Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reactive controls. MIKASA-Robo, the suite it rebuilds, has 32 tasks and uses language only in a representative VLA subset. Here every task provides an instruction, while memory-dependent tasks hide a task-relevant cue and reactive controls keep it available. For 70 tasks, environment phase timings specify an information gap, and for 28 of them the gap exceeds the 16-frame window of the widest fixed-context VLA we survey. The gap counts only the interval the cue is provably absent, not the full duration a policy must retain it, so every memory-dependent task still requires memory by construction, including the ones whose measured gap is short. We release 22,500 oracle trajectories across 10 memory types in RLDS and LeRobotDataset v3. A reference $π_{0.5}$ baseline with current images and proprioception, but no observation history or explicit memory module, is fine-tuned on 14 tasks and achieves 0.211 $\pm$ 0.044 mean task success. Its lower success on the evaluated Long-split tasks is confounded by open-loop chunking and the memory types represented in that subset. Project page: https://mikasarobo.github.io/