🤖 AI Summary
This study addresses the limitations of Vision-Language-Action (VLA) models, which suffer from long-term memory deficits due to constrained visual perception, as well as the poor interpretability and transferability of existing latent vector memories. To overcome these challenges, this work proposes AD-Memo, a novel agent framework that introduces a verbalized memory mechanism. This approach pioneers representing memories as Chain-of-Thought outputs that are fed back into future inputs. Furthermore, it presents Da Capo, a semi-closed-loop reinforcement learning algorithm designed to achieve precise credit assignment at both trajectory and step levels. Experimental results demonstrate that the proposed method significantly enhances autonomous driving quality and scene question-answering capabilities. Ultimately, it yields a highly interpretable, plug-and-play universal memory module suitable for broad deployment.
📝 Abstract
Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize knowledge acquired during vision-language pretraining for accurate and interpretable driving. However, VLAs can take only a limited number of frames as visual input due to the high token cost of an image, which is problematic for memory-dependent tasks such as determining the arrival order at all-way stops and long-horizon driving scene understanding. Existing solutions use latent vector memories accessed through cross-attention, which are neither interpretable nor portable. In this paper, we propose AD-Memo, a general-purpose VLA driving agent with language-based memory. The agent outputs memory as an extension of its Chain-of-Thought (CoT) to record surrounding objects critical to driving; this memory becomes part of the agent's future input. We curate memory-based datasets and train VLAs with a two-stage recipe: Supervised Fine-Tuning (SFT) and \textit{Da Capo}, a novel semi-closed-loop Reinforcement Learning (RL) algorithm which uses trajectory-level advantage for memory and step-level advantage for driving, leading to better credit assignment. Across scenarios such as all-way stops and general driving, AD-Memo improves driving quality, enables better question answering on driving scenes, and produces plug-and-play memory for other models.