TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing streaming video evaluation methods, which overlook evidence timeliness and triggering mechanisms, allowing similar scores to obscure fundamental differences in system behavior. To this end, we propose TRACE, a framework that explicitly quantifies information availability and event responsiveness through temporal auditing and condition-aware evaluation. Furthermore, it introduces a unified causal Core-Adapter protocol to enable multi-dimensional reporting of conditioned execution behaviors. Evaluations based on temporal annotations, causal controls, and a multi-dimensional metric suite reveal significant disparities in system workload and reliability despite identical accuracy. Consequently, this work establishes a new paradigm for streaming video evaluation.
📝 Abstract
Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report task scores without specifying when evidence becomes valid, how visual history is maintained, or how responses are triggered. As a result, similar scores may correspond to different workloads, failure modes, and operational behavior. We introduce TRACE (Temporal Audit and Condition-aware Evaluation), a condition-aware benchmark and evaluation framework that makes these factors explicit. TRACE combines temporally audited visual tasks with evidence timing and instruction-dependent trigger annotations, a unified causal Core--Adapter protocol that controls information availability while recording actual history processing and response events, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. On 1,240 records from 517 videos, we evaluate eight publicly available models or systems in eight configurations. We find that nearly identical QA accuracy can mask substantial differences in completion, answer validity, and generation workload, while proactive performance separates into response quality, response delay, false alarms (responses emitted while no target window is currently valid and a later one remains), and missed target windows. These results show that streaming-video performance should be interpreted as execution-conditioned system behavior rather than a single score. Our benchmark and code can be accessed at \href{https://github.com/om-ai-lab/trace-bench}{https://github.com/om-ai-lab/trace-bench}.
Problem

Research questions and friction points this paper is trying to address.

streaming video understanding
evaluation benchmark
temporal audit
condition-aware evaluation
response behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

streaming video understanding
condition-aware evaluation
temporal audit
causal Core-Adapter protocol
multidimensional reporting
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30