🤖 AI Summary
This study addresses the evaluation blind spots in existing long-term memory benchmarks caused by their neglect of multimodal implicit cues. To this end, we propose CUE-Mem, the first long-term memory evaluation benchmark focusing on implicit text-image-audio cues, encompassing four task categories including entity recall. By comparing textualized memory systems with native multimodal indexing techniques, this work quantifies the capability gap in models' processing of implicit cues, revealing that the retrieval and retention of subtle evidence constitute core bottlenecks. Experimental results demonstrate that augmenting descriptions incurs substantial computational costs, while native access is constrained by noise. Overall, this work establishes a novel testbed for evaluating multimodal long-term memory.
📝 Abstract
Long-term memory is essential for multimodal agents that interact with users across sustained conversations. However, user memories are not always explicitly stated: they may also be implied by recurring background objects in images, ambient sounds in audio, or other peripheral multimodal cues. Existing benchmarks largely focus on text-only memory or explicit multimodal evidence, leaving implicit multimodal cues underexplored. We introduce CUE-Mem, a text-image-audio benchmark for evaluating long-term user memory from implicit cues. CUE-Mem contains 2,674 questions across explicit and implicit evidence settings and covers four tasks: Entity Recall, Long Pattern, Personalized Recommendation, and Answer Refusal. Across textualized memory systems, implicit performance remains far below oracle evidence, locating the main bottleneck in preserving and retrieving subtle cues rather than question answerability. Increasing caption detail recovers more of this evidence, but brings uneven gains and rapidly growing token costs, motivating native multimodal access. Yet native access does not uniformly resolve the bottleneck: evidence use depends strongly on the backbone, while multimodal indexing introduces substantial retrieval noise. CUE-Mem provides a testbed for memory systems that selectively retain, retrieve, and use subtle multimodal evidence.