🤖 AI Summary
This work addresses the challenges of high redundancy, prohibitive memory costs from dense sampling, and the risk of missing critical evidence under sparse sampling in long video temporal localization. To this end, we propose a "summarize-then-localize" framework that introduces a novel query-guided latent summarization and associative retrieval mechanism to aggregate relevant evidence via chunk-wise compression. Additionally, a length-aware gradient gating module is incorporated to substantially reduce training overhead. The multimodal reasoning pipeline is further optimized through verifiable reward reinforcement learning, KV-state compression, and chunked processing techniques. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art approaches across multiple benchmark datasets, yielding particularly significant performance gains in long video scenarios.
📝 Abstract
Video temporal grounding (VTG) aims to localize the video interval corresponding to a language query. Recent large vision-language models (LVLMs) show great potential in solving such a multi-modal reasoning task. However, long videos often contain large amounts of redundant information that disturbs LVLMs to mine query-relevant evidence. Instead of dense frame sampling which incurs prohibitive training memory, previous reinforcement learning with verifiable rewards (RLVR) works typically utilize sparse sampling, which makes training feasible but may miss critical evidence. In this paper, we propose a ``summarize before grounding'' framework (named ``SumGround'') for long-video temporal grounding. The key of SumGround is to perform query-guided chunk condensation to aggregate and retrieve query-relevant evidence. Specifically, we split the video into several chunks and perform two-level chunk condensation. First, we introduce query-guided latent summaries, which is represented as KV states of query-guided prompts, to compress redundant visual tokens into compact query-relevant chunk summaries. Furthermore, we design an associative summary retrieval scheme to rank and select chunk summaries that are most likely to contain the event interval. Both query-guided latent summary and associative summary retrieval schemes are enabled by RLVR. To reduce memory consumption, we propose a length-aware gradient gating module to selectively stop gradient back-propagated to visual tokens. Extensive experiments demonstrate that SumGround performs favorably against previous state-of-the-art methods across multiple downstream datasets, with remarkable gains on long videos.