🤖 AI Summary
This study addresses the challenge of balancing recent and historical evidence under limited computational budgets in streaming video understanding by proposing WRWS, a training-free framework. Specifically, it quantifies retrieval uncertainty via the entropy of similarity distributions to filter keyframes prior to retrieval, dynamically adjusting the weights between semantic retrieval and recent priors for adaptive allocation of visual evidence. By integrating a lightweight external vision-language encoder with a retrieve-then-encode pipeline, WRWS functions as a plug-and-play module compatible with diverse models. Experimental results demonstrate that WRWS maintains competitive accuracy across multiple benchmarks while reducing the average visual-to-answer latency to 47.93% of that required by state-of-the-art methods.
📝 Abstract
Streaming video understanding requires Video Large Language Models (Video-LLMs) to reason over continuous visual streams under causal constraints. As the visual history grows, a bounded visual?processing budget requires evidence selection that balances temporal recency with query relevance. Recent-only selection excludes potentially relevant historical evidence, whereas Semantic-only retrieval can displace useful recent context when relevance scores are ambiguous. We introduce WRWS (When to Retrieve, When to Stay), a training-free framework for uncertainty-adaptive evidence allocation. A lightweight external vision-language encoder scores query relevance across the observed history, while an adaptive allocation module uses the normalized entropy of the similarity distribution as a proxy for retrieval uncertainty. WRWS favors semantic retrieval when relevance cues are reliable and strengthens the recency prior under uncertainty. Following a retrieve-first, encode-later pipeline, WRWS selects evidence before target-model visual encoding, such that only the selected observations are processed by the costly target Video-LLM. Experiments across four Video-LLM families and multiple model scales demonstrate competitive accuracy on StreamingBench and OVO-Bench. In our efficiency evaluation, WRWS reduces average vision-to-answer time to 47.93% of the state-of-the-art method. Code will be released.