Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance limitations of Vision-Language-Action (VLA) models in history-dependent tasks caused by the absence of effective memory mechanisms. We propose D&R, a recursive memory method grounded in the Partially Observable Markov Decision Process (POMDP) framework. For the first time, this work formulates memory selection as a conditional mutual information maximization problem. By designing a recursive divide-and-conquer architecture with a shared selector and a Top-K token selection mechanism, our approach overcomes fixed context window constraints, enabling efficient long-horizon reasoning using only 64 tokens. Experimental results demonstrate that the proposed method achieves state-of-the-art success rates across all 16 tasks on the RoboMME benchmark, and its effectiveness is further validated through real-world robot deployments.
📝 Abstract
Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pixel changes, and show inconsistent gains across tasks. We view what to remember as an optimisation problem. From the POMDP formulation of imitation learning, we show that the optimal memory maximises the conditional mutual information $I(a_t; m_t \mid o_t)$ between the action and the memory given the current observation. Intuitively, this means preserving the action-relevant information in the history that is not already contained in the current observation. Based on our analysis, we propose Divide-and-Remember (D&R), a recursive memory method that learns a memory function $m_t = M(h_t)$ and scales to long contexts while staying compute-light. It involves two strategies: (1) the selection over the full history is divided recursively into subproblems of top-$K$ selection over $2K$ tokens, so that fixed-size, lightweight selectors learned end-to-end support an unbounded history; (2) all recursion blocks share one selector, which captures the selection rule common to every block and keeps the method efficient. On RoboMME, a benchmark of 16 long-horizon manipulation tasks that require remembering when, where, what, and how to act, D&R achieves a state-of-the-art average success rate with consistent gains across all four suites under a budget of only 64 tokens; real-robot experiments show the same gain. Code, checkpoints and more results are at https://dnr-memory.github.io/
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
Long-horizon manipulation
History-dependent tasks
Memory optimization
Partially Observable Markov Decision Process
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
Recursive Memory
Conditional Mutual Information
Divide-and-Conquer
Long-Horizon Manipulation
🔎 Similar Papers
No similar papers found.