π€ AI Summary
This study addresses the challenges of designing fixed memory mechanisms and compressing unbounded visual streams in streaming video understanding by proposing an automated program search framework driven by large language models (LLMs). Specifically, we design a domain-specific language (DSL) to define memory primitives, thereby reformulating memory mechanisms as executable programs. An LLM is then leveraged to iteratively generate, verify, and optimize these programs, establishing a pioneering training-free automated research paradigm guided by empirical feedback. Extensive evaluations on benchmarks such as StreamingBench demonstrate that our approach achieves superior performance while significantly enhancing both contextual and reasoning efficiency. Furthermore, the proposed framework automatically discovers highly effective and bounded memory compression mechanisms, offering a scalable solution for continuous video understanding tasks.
π Abstract
Query-agnostic streaming video understanding requires vision-language models to continuously compress an indefinitely growing visual stream into a bounded memory before future queries are known. The performance depends critically on the memory mechanism--what observations to preserve, how to represent and consolidate them, and what information to retrieve when a query eventually arrives. Rather than designing a single memory architecture by hand, we formulate memory design as a search problem over executable memory programs. We introduce a lightweight domain-specific language that expresses memory mechanisms through structured primitives for representation, admission, retention, consolidation, budgeting, and retrieval, while enforcing causal and bounded-memory constraints. Although structured, the derived program space remains large and contains heterogeneous, conditionally dependent design choices whose effects can only be assessed via downstream execution. We therefore propose MemEvo, an LLM-driven auto-research framework that uses pretrained LLM as a semantics-aware proposal model to iteratively generate and refine candidate memory programs based on accumulated experimental feedback. At runtime, a deterministic evaluation pipeline validates and evaluates each candidate, while the underlying vision-language model remains frozen throughout discovery. We finally produce a training-free, bounded-memory mechanism. Extensive experiments on StreamingBench and OVO-Bench demonstrate strong streaming video understanding performance together with substantial context and inference efficiency.