Score
Designs, builds, and evaluates algorithms and models that extract and interpret spatiotemporal information from video data, including recognition, detection, segmentation, tracking, temporal localization, summarization, captioning, and retrieval of events and actions. Focuses on learning and testing representations that capture motion, appearance, and temporal structure to enable understanding, reasoning, and indexing of video content.
To address insufficient spatiotemporal modeling capabilities amid explosive growth in video content, this paper proposes a unified spatiotemporal analysis framework that jointly enhances long-range dependency modeling through multi-scale temporal modeling and dynamic spatial attention. Methodologically, the framework systematically integrates 3D CNNs, Transformer-based video encoders, and contrastive learning pretraining, and is comprehensively evaluated across multiple benchmarks—including Kinetics, Something-Something, and UCF101. The core contribution lies in the decoupled yet synergistic optimization of spatial and temporal representations: dynamic attention adaptively focuses on salient frames and regions, while multi-scale temporal modules hierarchically capture both short-term motion patterns and long-term semantic structures. Experiments demonstrate an average 4.2% improvement in action recognition accuracy, an 18% reduction in temporal localization error, and concurrent gains in model generalizability and interpretability.
Current Video-LLMs exhibit fundamental limitations in modeling abstract temporal concepts—such as long-range dependencies, causal relationships, and event evolution—due to the absence of explicit temporal annotations in video datasets, domain-specific biases, and a temporal misalignment between visual encoders and LLMs. Method: We systematically diagnose performance gaps in cross-segment event association and causal reasoning; propose a novel spatio-temporal semantic joint modeling paradigm; and construct a multi-source video data framework with explicit temporal annotations. We further conduct temporal attribution analysis, multimodal fusion modeling, and bias assessment to validate our approach. Contribution/Results: Our method significantly enhances temporal awareness in Video-LLMs, demonstrating measurable improvements in temporal reasoning tasks. The framework is fully reproducible and empirically verifiable, offering a principled technical pathway toward next-generation Video-LLMs with robust, interpretable temporal understanding capabilities.
Real-time video analysis faces a fundamental trade-off between spatiotemporal modeling accuracy and inference efficiency, particularly under resource constraints. To address this, we propose a unified multi-task framework that jointly performs action recognition and object tracking. Our approach employs a hierarchical attention mechanism to adaptively focus on temporally salient spatial regions and introduces parallel sequence modeling to enhance computational efficiency. By integrating advanced spatiotemporal representation learning with lightweight co-optimization strategies, the method achieves state-of-the-art performance: +3.2% and +2.8% top-1 accuracy on UCF-101 and HMDB-51 for action recognition, +2.8% MOTA on MOT17 for multi-object tracking, and a 40% speedup in end-to-end inference latency—significantly outperforming existing real-time approaches.
Addressing the challenge of jointly modeling fine-grained (pixel-level) and coarse-grained (regional-level) spatial structures while preserving long-range temporal dependencies in traffic video forecasting, this paper proposes a tile-based spatial attention-enhanced encoder-decoder LSTM framework. Our key contributions are: (1) a tile-aware spatial attention mechanism that explicitly captures multi-scale spatial hierarchies; (2) an attention-guided LSTM cell integrating convolutional features to improve trajectory modeling fidelity; and (3) an adaptive frame-sampling strategy coupled with an end-to-end trainable encoder-decoder architecture, balancing computational efficiency and robustness. Evaluated on the Traffic4Cast 2019 dataset, our method significantly outperforms both 2D/3D CNNs and standard ConvLSTM baselines, achieving superior prediction accuracy and enhanced spatiotemporal consistency.
This work addresses the challenge of high annotation costs and the availability of only video-level weak labels in video anomaly detection by proposing a novel weakly supervised approach that, for the first time, jointly models spatiotemporal local anomalies under such constraints. The method treats normal and anomalous videos as negative and positive bags, respectively, within a multiple instance learning (MIL) framework. It integrates spatiotemporal feature extraction with a classifier-driven anomaly scoring mechanism and introduces a multiple instance ranking loss to effectively leverage video-level labels for pixel-level anomaly localization. Experimental results on the UCF Crime2Local dataset demonstrate that the proposed method accurately detects localized spatiotemporal anomalies using only video-level supervision.
This work proposes treating temporal flow rate as a learnable visual concept to enable perception and controllable generation of video playback speed. Through a self-supervised approach leveraging multimodal cues and temporal structure inherent in videos, the model accurately detects speed variations and estimates actual playback rates. Building upon this framework, the authors construct the largest wild slow-motion video dataset to date and demonstrate two key applications: synthesizing videos at user-specified playback speeds and performing temporal super-resolution on low-frame-rate videos to recover fine-grained dynamic details. This study is the first to model time flow rate as a manipulable perceptual dimension, significantly advancing capabilities in video temporal understanding and generation.
Existing research on video event detection lacks a unified large-scale dataset and standardized evaluation protocols, hindering fair method comparison and reproducibility. To address this gap, this work proposes the first integrated, three-pronged development framework encompassing dataset construction, performance evaluation, and deployment scenarios. By introducing structured data design, a standardized metric system, and diverse application-oriented modeling, the framework establishes a generalizable paradigm for the field. This approach substantially enhances the fairness of algorithmic comparisons, improves research reproducibility, and supports systematic methodological advancement in video event detection.
This work addresses the limitation of existing video moment retrieval and highlight detection methods, which often overlook fine-grained intra-frame visual details guided by text, thereby constraining localization accuracy. To overcome this, the authors propose a unified spatiotemporal representation learning framework that, for the first time, jointly models text-driven progressive fine-grained image encoding and multi-scale temporal dynamics within an end-to-end architecture. By enhancing the collaborative alignment between frame-level spatial details and temporal context, the approach substantially improves video-text semantic matching. The method achieves state-of-the-art performance across four benchmark datasets: QVHighlights, Charades-STA, TACoS, and TVSum.