Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action Policies

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the perceptual aliasing and low success rates in long-horizon tasks exhibited by Vision-Language-Action (VLA) models due to the absence of interaction history. To this end, we propose ActMem-VLA, an architecture that leverages a Mamba state space model to encode action memory. We introduce a novel dual-expert handoff mechanism under a frozen backbone, wherein a lightweight pre-action expert performs memory-guided steering while a diffusion denoising module executes fine-grained control. This design decouples history-conditioned planning from action refinement, substantially reducing training costs. On the LIBERO-Mem benchmark, ActMem-VLA achieves an average success rate of 80.8% with only a 3.45% parameter increase. Furthermore, it outperforms π0.5 by 28.8% on real-world tasks.
📝 Abstract
Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current observations and robot states at different task stages may induce action ambiguity and lower success rate. Existing methods incorporate temporal or progress cues through feature conditioning, action-prior modification, or sampling guidance. However, methods that jointly fine-tune memory modules and the base VLA incur additional policy-training costs, motivating the separation of trainable history-conditioned steering from frozen base-policy refinement. We propose ActMem-VLA, a dual-expert handover architecture that augments a frozen, fine-tuned VLA with a memory plugin comprising a Mamba-based memory module and a lightweight PreAction Expert (PAE). Specifically, Mamba encodes executed-action history into memory that conditions PAE alongside current context. With these inputs, PAE steers task progression during early, high-noise denoising, then passes the partially denoised action to the frozen Action Expert (AE) to refine action details during the remaining low-noise steps. The fine-tuned base VLA remains frozen throughout training, while only the Mamba module and PAE are jointly optimized. On LIBERO-Mem, ActMem-VLA achieves 80.8\% average success across all ten tasks, compared with 65.2\% for $π_{0.5}$ and 49.5\% for MemoryVLA, while introducing only 3.45\% additional parameters. Across four real-world tasks, it improves the average success rate over $π_{0.5}$ by 28.8\%.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
Perceptual aliasing
Action ambiguity
Long-horizon tasks
Robotic manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Dual-Expert Denoising
Mamba Memory Module
Perceptual Aliasing
Parameter-Efficient Fine-Tuning
💼 Related Jobs
No related jobs found.
Y
Yaxin Zhao
Harbin Institute of Technology, Harbin, China; Medical Intelligence and Robotic Cognition (MIRoC) Lab, Department of Mechanical Engineering, The University of Hong Kong (HKU), Hong Kong SAR, China
Dianye Huang
Dianye Huang
Technical University of Munich
robotic ultrasoundmedical robotintelligent controlhuman robot interaction
C
Chenwei Wang
Medical Intelligence and Robotic Cognition (MIRoC) Lab, Department of Mechanical Engineering, The University of Hong Kong (HKU), Hong Kong SAR, China
Chenguang Yang
Chenguang Yang
Chair Professor in Robotics, Fellow of IEEE, IET, IMechE, AIAA, BCS
Robotics
Zhongliang Jiang
Zhongliang Jiang
University of Hong Kong
Medical RoboticsUltrasound imagingRobot learningSurgical RoboticsHuman-robot Interaction