🤖 AI Summary
This study addresses the challenge faced by existing vision-language models in streaming video understanding, where balancing low latency with long-context retention often leads to progressive information loss as inputs grow. To this end, this work proposes ReMem, a training-free adaptation framework that leverages a dual-memory mechanism to process streaming videos of arbitrary length. Specifically, it introduces Streaming Context Memory (SCM) for query-agnostic attention compression and Retrieval Visual Memory (RVM) to extract salient, relevant context for input augmentation. By synergistically modeling compression and retrieval, ReMem effectively resolves the core issue of progressive context degradation in long-sequence scenarios. Experiments demonstrate that ReMem achieves state-of-the-art performance across multiple mainstream benchmarks, significantly enhancing both streaming and general long-form video understanding capabilities.
📝 Abstract
Despite their impressive performance on a wide range of video understanding tasks, current Vision Language Models (VLMs) are predominantly designed for offline scenarios and struggle to handle online streaming videos that demand low latency response. Several studies have explored memory and token compression strategies in an attempt to adapt offline VLMs for streaming video understanding tasks. However, through our probing experiment, we identify that most existing works tend to progressively lose long context information as length of input stream increases. To address this, we propose ReMem, a novel training-free adaptation technique that enables VLMs to process streaming videos of arbitrary lengths while improving their long context information retention capability. ReMem exploits memory from two perspectives, implemented as two core components. The Streaming Context Memory (SCM) continuously compresses historical context with query-independent attention. The Retrieved Vision Memory (RVM) then retrieves the most salient, query-relevant context from memory to augment the VLM's input. Comprehensive experiments demonstrate that the proposed ReMem achieves state-of-the-art (SOTA) performance across a variety of widely used benchmarks, spanning both streaming video and general long video understanding tasks.