STORM-Bench: Evaluating Online Video QA under Evolving and Incomplete Evidence

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing benchmarks in evaluating state tracking and refusal capabilities for video question answering under conditions of dynamically evolving and incomplete evidence. To this end, we construct a dynamic video QA benchmark comprising 5,736 questions drawn from both real-world first-person scenarios and simulated environments. We further propose the STORM-BR metric to jointly assess answer correctness and refusal reliability, and conduct a systematic evaluation of fourteen video large language models using a 1 FPS hierarchical sampling strategy. Experimental results reveal that mainstream models achieve an average accuracy of only 51.7%, while the reliability metric drops to 18.8%. These findings expose critical overconfidence and cognitive reliability deficiencies that remain obscured by conventional task accuracy measures.
📝 Abstract
Reliable online video question answering requires tracking state transitions while selectively abstaining when visual evidence is insufficient. Existing benchmarks focus on static recognition or long-range retrieval, rarely evaluating these coupled capabilities under evolving and incomplete evidence. We present STORM-Bench, comprising 5,736 questions across 630 compact, change-dense episodes spanning five egocentric domains (STORM-Real) and two controlled simulation subsets (STORM-Sim) at 1 FPS. Questions are stratified by a proxy for accumulated change intensity (Low, Medium, High) and query-time answerability (Known, Uncertain). To measure reliability, we introduce STORM-BR, a harmonic metric over joint answer-status correctness that exposes abstention failures masked by aggregate accuracy, alongside STORM-BR-ATTR for uncertainty attribution. Across 14 video LLMs, online accuracy peaks at 60.3\% (mean 51.7\%), whereas STORM-BR ranges from 5.7\% to 35.6\% (mean 18.8\%), driven by pervasive overconfidence on uncertain queries. STORM-Bench shows that task accuracy masks these gaps in epistemic reliability and state tracking. Benchmark and code are available at https://github.com/siruzhong/STORM-Bench.
Problem

Research questions and friction points this paper is trying to address.

Online Video QA
State Tracking
Epistemic Reliability
Incomplete Evidence
Benchmark Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Online Video QA
STORM-Bench
Epistemic Reliability
Selective Abstention
State Tracking
🔎 Similar Papers
2024-08-08International Journal of Computer VisionCitations: 13
2024-10-10arXiv.orgCitations: 0
2024-02-20International Conference on Machine LearningCitations: 30