TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes

📅 2025-02-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing video understanding models lack unified spatiotemporal reasoning capabilities for complex, real-world roadside traffic scenarios. Method: We introduce TrafficBench—the first multi-task benchmark for realistic traffic understanding—comprising 1,000 videos, 85K multiple-choice QA pairs, 2.3K referring object descriptions, and 5.7K spatiotemporal object localization annotations, covering challenging conditions such as adverse weather and anomalous events. We propose a novel tuple-based spatiotemporal object representation paradigm that unifies multiple-choice video QA, referring object captioning, and spatiotemporal object localization within a single evaluation framework. Additionally, we design a visual token sampling strategy to enhance the Qwen architecture for fine-grained spatiotemporal reasoning. Results: Extensive experiments expose significant limitations of current foundation models on these tasks, establishing TrafficBench as a high-difficulty, open-source, and reproducible benchmark to advance intelligent transportation systems.

Technology Category

Knowledge Representation and Reasoning: Geometric, Spatial, and Temporal ReasoningComputer Vision: Visual Reasoning & Symbolic RepresentationsPlanning, Routing, and Scheduling: Optimization of Spatio-temporal Systems

Application Category

Search and Retrieval-Augmented AI: Web query analysis, representation and understandingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSystems and Infrastructure for Web, Mobile and WoT: Web performance, measurement, and characterization
📝 Abstract
We present TUMTraffic-VideoQA, a novel dataset and benchmark designed for spatio-temporal video understanding in complex roadside traffic scenarios. The dataset comprises 1,000 videos, featuring 85,000 multiple-choice QA pairs, 2,300 object captioning, and 5,700 object grounding annotations, encompassing diverse real-world conditions such as adverse weather and traffic anomalies. By incorporating tuple-based spatio-temporal object expressions, TUMTraffic-VideoQA unifies three essential tasks-multiple-choice video question answering, referred object captioning, and spatio-temporal object grounding-within a cohesive evaluation framework. We further introduce the TUMTraffic-Qwen baseline model, enhanced with visual token sampling strategies, providing valuable insights into the challenges of fine-grained spatio-temporal reasoning. Extensive experiments demonstrate the dataset's complexity, highlight the limitations of existing models, and position TUMTraffic-VideoQA as a robust foundation for advancing research in intelligent transportation systems. The dataset and benchmark are publicly available to facilitate further exploration.
Problem

Research questions and friction points this paper is trying to address.

Spatio-temporal video understanding in traffic
Unifying multiple video analysis tasks
Enhancing intelligent transportation systems research
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatio-temporal object expressions
Visual token sampling strategies
Unified evaluation framework
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30