VideoTapestry: Query-Adaptive Memory Refinement for Multi-Agent Long-Video Understanding

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the susceptibility of memory retrieval to localization errors and the omission of fine-grained details in long video understanding. To tackle these challenges, we propose a training-free multi-agent collaborative framework that introduces a novel hierarchical video memory architecture. By employing a coarse-to-fine query-driven strategy, the method dynamically optimizes pre-constructed memories, enabling adaptive reconstruction that preserves global context while enhancing fine-grained evidence. Extensive evaluations on benchmarks such as LVBench demonstrate that the proposed framework improves accuracy by up to 17.2%, establishing new state-of-the-art performance across multiple long video understanding tasks.
📝 Abstract
Long-video understanding places substantial demands on memory, as answering questions often requires retrieving information distributed across extended temporal spans. Existing approaches broadly follow two paradigms: query-driven exploration, which is sensitive to localization errors, and query-independent memory construction, which may omit question-specific details. We introduce VideoTapestry, a training-free multi-agent framework that adapts a preconstructed hierarchical video memory through coarse-to-fine, query-driven refinement. The preconstructed memory organizes video content into three levels, capturing global narrative context, event-level temporal structure, and fine-grained relational evidence, respectively. To support coarse-to-fine localization and observation, we assign a specialized agent to each level, keeping retrieval and refinement within a scale-specific context. Guided by the query, these agents revisit relevant video regions and enrich layer-wise memories with targeted multimodal observations. Their refinements are assembled according to the original hierarchy into a composite query-adaptive memory, preserving global context in a compact form while retaining fine-grained evidence along query-relevant branches for final reasoning. Compared with direct GPT-5.5 inference, VideoTapestry achieves absolute accuracy gains of 17.2%, 14.9%, 9.8%, and 7.0% on LVBench, LongVideoBench (Long), Video-MME (Long), and EgoSchema, respectively, achieving the state-of-the-art results among all competitors.
Problem

Research questions and friction points this paper is trying to address.

Long-video understanding
Memory refinement
Multi-agent framework
Query-adaptive retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Framework
Long-Video Understanding
Query-Adaptive Memory
Hierarchical Memory Refinement
Training-Free
🔎 Similar Papers
fetch failed