DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing multimodal benchmarks in neglecting the evaluation of dense, dynamic visual state memory within professional workflows. To this end, it introduces the concept of dense stateful visual memory alongside a Hartley-inspired heuristic criterion. Methodologically, this work pioneers a decoupled generation framework that separates state synthesis from dialogue population, constructs an expert-verified benchmark encompassing five query categories, and designs automated generation with fine-grained evaluation systems. Experimental results demonstrate that state-of-the-art models score below 45%, identifying state update frequency as the primary bottleneck. Furthermore, findings reveal that state-aware design significantly outperforms merely scaling inference compute.
📝 Abstract
Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios featuring photographic natural images, isolated static artifacts, and recall-oriented questions. In contrast, professional scenarios often involve structured, information-heavy artifacts that undergo frequent revisions and authority updates, and compositional queries requiring reconciliation of many artifact versions while tracking state precisely. To address these challenges, we introduce DSV-Mem, a benchmark for evaluating Dense Stateful Visual Memory. DSV-Mem comprises expert-reviewed scenarios and 1,000 questions across five user-oriented categories (Current State, Past State, Derived State, Change History, and Conflict/Refusal). A Hartley-inspired criterion favors questions with broader visual-evidence inspection demands. We also introduce a generation harness that produces evaluation suites by decoupling state-transition synthesis from conversation filling. Evaluation over 27 configurations spanning frontier and open-weight models and memory management methods reveals that the strongest baseline scores below 45% on DSV-Mem. Analysis surfaces findings: 1) multimodality and information density both contribute to difficulty, but state evolution, particularly the number of governing updates, is the dominant tested factor. Raw conversation/haystack length, OCR, and arithmetic are not the primary bottlenecks; 2) models often fail to verify user premises against prior state updates before answering; 3) increased reasoning effort and memory management methods yield limited gains, whereas state-aware designs prove more effective. The benchmark and code will be publicly released.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Memory
MLLM Agents
Professional Workflows
Benchmark Evaluation
Stateful Visual Artifacts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dense Stateful Visual Memory
Multimodal Benchmark
MLLM Agents
State-Transition Synthesis
Professional Workflows
🔎 Similar Papers
2024-10-04International Conference on Learning RepresentationsCitations: 0
💼 Related Jobs
No related jobs found.
Jike Zhong
Jike Zhong
University of Southern California
Computer VisionMachine Learning
Ritwick Chaudhry
Ritwick Chaudhry
Amazon Web Services
Computer VisionMachine LearningDeep Learning
X
Xuanbai Chen
Amazon AGI
T
Tianchen Zhao
Amazon AGI
L
Linghan Xu
Amazon AGI
Y
Yifan Xing
Amazon AGI
N
Nishant Sankaran
Amazon AGI