PMMC: Prospective Multimodal Memory Compilation for Long-Term LVLM Agents

๐Ÿ“… 2026-08-01
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the inefficiency and poor robustness of existing large vision-language model (LVLM) agents in handling queries requiring image-text binding, temporal updates, or fine-grained visual information due to limitations in their long-term memory systems. To overcome this, the authors propose a forward-looking multimodal memory compilation framework that anticipates potential queries during the memory consolidation phase, compiles conditional memory programs, and validates evidence pathways through a questioning mechanism. Introducing the novel โ€œmemory-as-computationโ€ paradigm, the framework employs a synergistic architecture integrating a question generator, planner, and skeptic to construct a structured question bank for efficient retrieval. Experiments demonstrate that this approach significantly improves answer quality and visual evidence recall on multimodal long-term memory benchmarks while reducing token consumption and latency at query time.
๐Ÿ“ Abstract
Long-term memory is essential for LVLM agents to maintain consistency and integrate information across extended multimodal interactions. Existing agent memory systems, however, often reduce visual experiences into textual summaries or rely on static retrieve-then-reason pipelines, which are inefficient at query time and brittle when questions require image-text binding, temporal updates, or visual details. We propose Prospective Multimodal Memory Compilation, a framework that shifts part of the memory reasoning process from query time to memory consolidation time. Given accumulated multimodal interactions, a Questioner predicts future question candidates, a Planner compiles question-conditioned multimodal memory programs, and a Doubter verifies whether the planned evidence path can support the predicted answer. The verified question-program pairs form a structured question bank for efficient query-time routing and evidence retrieval. Experiments on multimodal long-term memory benchmarks show that our method improves answer quality and visual evidence recall while reducing query-time token and latency costs. Extensive ablations analyze the effects of self-feedback, dynamic planning, raw-image access, and question bank coverage.
Problem

Research questions and friction points this paper is trying to address.

long-term memory
multimodal interaction
image-text binding
query efficiency
visual detail
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prospective Memory Compilation
Multimodal Long-Term Memory
Query-Time Efficiency
Memory Consolidation
LVLM Agents
๐Ÿ”Ž Similar Papers
2024-10-04International Conference on Learning RepresentationsCitations: 0
Jingyu Sun
Jingyu Sun
NTT Corporation
Deep learningOntologyFew shot learningTime seriesActive learning
Y
Yan Lin
The University of Newcastle
Yuyang Xue
Yuyang Xue
PhD Student, University of Edinburgh
Machine UnlearningMRI ReconstructionComputer VisionRobustness
Y
Yifan Wang
The University of Manchester
Z
Zhengtao Yao
The University of Southern California
R
Rui Qian
Fudan University
Z
Zefeng Xu
The University of Manchester
J
Jiachen Li
The University of Texas at Austin
X
Xianyang Liu
Independent Researcher
Jiancheng Pan
Jiancheng Pan
INSAIT, Sofia University "St. Kliment Ohridski"; R.A., THU; M.S., ZJUT
Multimodal LearningFoundation ModelsMLLMsData-centric AIAI4Earth
Jingyuan Sun
Jingyuan Sun
Assistant Professor, The University of Manchester
neural encoding and decodingbrain machine interfacelarge language models
S
Syed Murtuza Baker
The University of Manchester
H
Hongpeng Zhou
The University of Manchester