TAU-Bench: From Anomaly Instance Tracking to Fine-Grained Video Anomaly Understanding

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical gap in video anomaly understanding: the disconnect between semantic description and spatiotemporal coherence of anomalous instances, along with the absence of a unified benchmark for jointly evaluating instance tracking and fine-grained semantic interpretation. To bridge this gap, we propose the first trajectory-centric evaluation framework that integrates pixel-level masks, event-level annotations, and scene-level captions to holistically assess a model’s ability to both track anomalous instances and provide accurate semantic explanations. We introduce an automated data construction pipeline—incorporating anomaly suitability filtering, trajectory generation, hierarchical captioning, and human quality control—to build a large-scale benchmark comprising 1,118 videos and 1,454 annotated trajectories. Empirical evaluation reveals that current vision-language models often produce fluent yet semantically misaligned descriptions due to inaccurate instance grounding, underscoring the fundamental challenge of visual grounding in anomaly understanding.
📝 Abstract
Humans understand anomalous events through a coherent perceptual process in which they identify the focal instance, follow its behavior as the event unfolds, and interpret why it violates the expectations of the surrounding scene. Video anomaly understanding (VAU) seeks to endow models with a similar capability, moving beyond deciding whether a video is anomalous toward explaining how the event develops and why it matters. Although recent vision--language models (VLMs) can generate detailed and plausible anomaly descriptions, their semantic fluency does not ensure that these interpretations remain grounded in the correct anomaly instance over time. Existing benchmarks typically evaluate tracking and semantic understanding through separate protocols, leaving such instance--semantic inconsistency largely unmeasured. We therefore introduce TAU-Bench, a track-centric benchmark for jointly evaluating anomaly instance tracking and fine-grained anomaly understanding. TAU-Bench contains 1,118 videos, 1,454 tracks, and 202,438 pixel-level masks spanning 49 event and 45 scene categories, together with track-centric annotations that connect instance-level identification, event-level understanding, and scene-level reasoning. To build TAU-Bench at scale, we developed an automated data engine integrating anomaly suitability filtering, anomaly instance track construction, hierarchical caption annotation, and human quality control. Evaluations across representative VLM families show that models producing plausible anomaly interpretations may still fail to localize and track the correct instance reliably, revealing a persistent gap between semantic reasoning and visual grounding. These findings therefore highlight instance-grounded evaluation as an important step toward more faithful and reliable VAU systems.
Problem

Research questions and friction points this paper is trying to address.

video anomaly understanding
anomaly instance tracking
instance-semantic consistency
visual grounding
fine-grained understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

video anomaly understanding
anomaly instance tracking
vision-language models
instance-grounded evaluation
TAU-Bench
🔎 Similar Papers