When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of balancing recent and historical evidence under limited computational budgets in streaming video understanding by proposing WRWS, a training-free framework. Specifically, it quantifies retrieval uncertainty via the entropy of similarity distributions to filter keyframes prior to retrieval, dynamically adjusting the weights between semantic retrieval and recent priors for adaptive allocation of visual evidence. By integrating a lightweight external vision-language encoder with a retrieve-then-encode pipeline, WRWS functions as a plug-and-play module compatible with diverse models. Experimental results demonstrate that WRWS maintains competitive accuracy across multiple benchmarks while reducing the average visual-to-answer latency to 47.93% of that required by state-of-the-art methods.
📝 Abstract
Streaming video understanding requires Video Large Language Models (Video-LLMs) to reason over continuous visual streams under causal constraints. As the visual history grows, a bounded visual?processing budget requires evidence selection that balances temporal recency with query relevance. Recent-only selection excludes potentially relevant historical evidence, whereas Semantic-only retrieval can displace useful recent context when relevance scores are ambiguous. We introduce WRWS (When to Retrieve, When to Stay), a training-free framework for uncertainty-adaptive evidence allocation. A lightweight external vision-language encoder scores query relevance across the observed history, while an adaptive allocation module uses the normalized entropy of the similarity distribution as a proxy for retrieval uncertainty. WRWS favors semantic retrieval when relevance cues are reliable and strengthens the recency prior under uncertainty. Following a retrieve-first, encode-later pipeline, WRWS selects evidence before target-model visual encoding, such that only the selected observations are processed by the costly target Video-LLM. Experiments across four Video-LLM families and multiple model scales demonstrate competitive accuracy on StreamingBench and OVO-Bench. In our efficiency evaluation, WRWS reduces average vision-to-answer time to 47.93% of the state-of-the-art method. Code will be released.
Problem

Research questions and friction points this paper is trying to address.

Streaming Video Understanding
Video-LLMs
Evidence Selection
Temporal Recency
Query Relevance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uncertainty-Aware Retrieval
Streaming Video-LLMs
Training-Free Framework
Temporal Evidence Allocation
Retrieve-First Pipeline
🔎 Similar Papers
X
Xiang Hu
IIAU Lab, Dalian University of Technology
Jiazuo Yu
Jiazuo Yu
Dalian University of Technology
Continual learningMulti-modal large language models
L
Lu Zhang
IIAU Lab, Dalian University of Technology
Yunzhi Zhuge
Yunzhi Zhuge
Dalian University of Technology
Computer Vision
H
Huchuan Lu
IIAU Lab, Dalian University of Technology