🤖 AI Summary
This study addresses the perceptual aliasing and low success rates in long-horizon tasks exhibited by Vision-Language-Action (VLA) models due to the absence of interaction history. To this end, we propose ActMem-VLA, an architecture that leverages a Mamba state space model to encode action memory. We introduce a novel dual-expert handoff mechanism under a frozen backbone, wherein a lightweight pre-action expert performs memory-guided steering while a diffusion denoising module executes fine-grained control. This design decouples history-conditioned planning from action refinement, substantially reducing training costs. On the LIBERO-Mem benchmark, ActMem-VLA achieves an average success rate of 80.8% with only a 3.45% parameter increase. Furthermore, it outperforms π0.5 by 28.8% on real-world tasks.
📝 Abstract
Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current observations and robot states at different task stages may induce action ambiguity and lower success rate. Existing methods incorporate temporal or progress cues through feature conditioning, action-prior modification, or sampling guidance. However, methods that jointly fine-tune memory modules and the base VLA incur additional policy-training costs, motivating the separation of trainable history-conditioned steering from frozen base-policy refinement. We propose ActMem-VLA, a dual-expert handover architecture that augments a frozen, fine-tuned VLA with a memory plugin comprising a Mamba-based memory module and a lightweight PreAction Expert (PAE). Specifically, Mamba encodes executed-action history into memory that conditions PAE alongside current context. With these inputs, PAE steers task progression during early, high-noise denoising, then passes the partially denoised action to the frozen Action Expert (AE) to refine action details during the remaining low-noise steps. The fine-tuned base VLA remains frozen throughout training, while only the Mamba module and PAE are jointly optimized. On LIBERO-Mem, ActMem-VLA achieves 80.8\% average success across all ten tasks, compared with 65.2\% for $π_{0.5}$ and 49.5\% for MemoryVLA, while introducing only 3.45\% additional parameters. Across four real-world tasks, it improves the average success rate over $π_{0.5}$ by 28.8\%.