🤖 AI Summary
This study addresses a critical vulnerability in existing spatio-temporal video grounding models, which continue to generate predictions even when presented with irrelevant queries, revealing their insufficient reliance on textual semantics. Through systematic ablation studies and dataset distribution analyses, we uncover statistical biases in benchmarks such as HCSTVG-v2 that incentivize models to disregard textual information, demonstrating that state-of-the-art approaches fail to genuinely comprehend image-text correlations. Furthermore, this work introduces a novel evaluation perspective centered on negative query awareness to rigorously validate these query-agnostic vulnerabilities. Ultimately, we advocate for the development of new architectures and evaluation protocols that explicitly assess query relevance, thereby promoting more robust and semantically grounded video understanding systems.
📝 Abstract
Spatio-temporal video grounding (STVG) aims to localize objects or events described by natural language queries in both space and time. Existing STVG models are typically trained and evaluated under the assumption that each query is relevant to the input video. In this work, we challenge this assumption by studying the behavior of state-of-the-art STVG models under irrelevant queries and missing textual input. Our experiments show that current models can still produce plausible spatio-temporal predictions even when the query is unrelated to the video or removed entirely. We further analyze HCSTVG-v2 and VidSTG to identify dataset regularities that may encourage such query-insensitive behavior. Our study highlights an underexplored limitation of STVG models and motivates negative-aware evaluation protocols and architectures that explicitly assess query relevance.