🤖 AI Summary
This study addresses the inability of existing video agent evaluation benchmarks to pinpoint hallucination sources and the disconnect between stage-wise scoring and downstream performance. We propose a causal stage intervention protocol that quantifies the genuine impact of specific modules on downstream tasks by systematically masking them. Through multi-architecture comparisons and large-scale experiments, we reveal that temporal grounding constitutes the dominant source of error, while region-level localization outperforms precise temporal overlap. Furthermore, erroneous evidence proves more detrimental than missing information, and conventional metrics are shown to be misleading. This work demonstrates the unreliability of prevailing scoring paradigms and advocates for a novel evaluation framework centered on causal intervention.
📝 Abstract
Video understanding is increasingly performed by multi-stage LLM agents that separate temporal grounding, visual observation, and reasoning. Yet these stages are typically evaluated on different benchmarks and distributions, making it difficult to determine where hallucinations originate. We first organize existing benchmarks around these stages and show that their scores provide inconsistent diagnostic signals: stronger stage-level performance does not reliably imply lower downstream hallucination, and even benchmarks targeting the same capability can disagree.
We therefore introduce a causal stage-intervention protocol that overwrites individual stages while holding the downstream task fixed. Across 60,008 runs on three video-agent architectures, we find that grounding is the dominant source of downstream error, with roughly four times the causal impact of corrupting visual observations. Successful grounding depends primarily on locating the correct region rather than precise temporal overlap, explaining why standard mIoU metrics poorly predict downstream reliability. We further find that incorrect evidence is substantially more harmful than missing evidence. Finally, auditing existing benchmarks against these interventions reveals that their scores do not reliably predict causal cascade sensitivity and can fail under distribution shift. These results motivate intervention-based, stage-aware evaluation for trustworthy video agents.