VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the semantic oscillation and loss of critical evidence in long-video agents caused by append-only memory growth. To this end, we propose a dual-loop multimodal agent architecture that introduces a novel inner-outer coupled bounded memory rewriting mechanism. Specifically, the inner loop dynamically updates working memory via file system retrieval, effectively mitigating context inflation and noise accumulation while preserving the validity of memory representations. Designed as a plug-and-play module, our method achieves an average improvement of 4.2% across four large vision-language model (LVLM) baselines. Furthermore, it attains a blind evaluation accuracy of 81.1%β€”peaking at 88.3%β€”on challenging questions, significantly enhancing the robustness of long-video reasoning.
πŸ“ Abstract
Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from semantic thrashing: as append-only working memory grows, attention to key evidence collapses, and the agent loses access to what it has already found. First, we provide a structural argument showing that append-only memory can incorporate newly observed target evidence, but cannot remove accumulated noise or prevent ordered context growth without a rewrite operator. Second, motivated by this analysis, we propose VideoLoop, a multimodal agent with two coupled loops. The outer loop reasons over the video and the inner loop, after each step, retrieves artifacts from an unbounded filesystem of past observations and intermediate analysis, and rewrites a bounded working memory. Extensive experiments demonstrate the effectiveness of VideoLoop, which improves four popular LVLM backbones in a plug-and-play manner, with an average gain of 4.2% points over baseline on VideoMME (long). Further analysis of working memory suggests that VideoLoop mitigates semantic thrashing: on the hardest quarter of VideoMME (long) questions, a blind judge that reads only the agent's context answers 81.1% correctly, versus 60.9% for the append-only agent. With Gemini 3.1 Pro, VideoLoop reaches 88.3% on VideoMME (long), 88.8% on VideoMMMU, and 80.9% on LongVideoBench (long).
Problem

Research questions and friction points this paper is trying to address.

Long-form video understanding
Semantic thrashing
Multimodal agents
Working memory
Append-only memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Thrashing
Looped Working Memory
Long-Form Video Understanding
Multimodal Agents
Memory Rewriting
πŸ”Ž Similar Papers