Score
Designs and implements algorithms that automatically identify and localize spatial regions of interest in video sequences by analyzing temporal changes (e.g., frame-to-frame or pixel variance) to detect activity or salient areas. Builds modules that output temporally consistent ROI masks or bounding boxes and perform cropping/masking to reduce background and camera effects.
Addressing the challenges of modeling spatial context, reliance on pre-trained models, and poor interpretability in street-scene video anomaly detection, this paper proposes an unsupervised spatial consistency modeling framework. First, Gaussian Mixture Modeling (GMM) is applied to high-resolution feature maps for unsupervised clustering, jointly discovering object-level spatial attributes and spatially consistent regions. Subsequently, an inter-object spatial relation graph is constructed to generate pixel-level normality heatmaps. The method requires neither pre-trained segmentation models nor human annotations. Evaluated on the Street Scene dataset, it achieves state-of-the-art performance while reducing parameter count by one to two orders of magnitude. Moreover, it produces high-resolution, semantically interpretable anomaly localization maps—enhancing model transparency and enabling efficient deployment.
Video foundational analysis requires efficient shot boundary detection, sampling pattern identification, and dynamic keyframe extraction. This paper proposes the first unified framework that simultaneously addresses hard-cut and short gradual shot boundary detection, discriminates scan-based sampling modes (progressive, interlaced, and film stretch), and performs motion-adaptive keyframe selection. Our method innovatively integrates motion field estimation with normalized cross-correlation (NCC) features, jointly models inter-frame and intra-frame multi-scale features, and introduces a sparse selective computation strategy—achieving 4× real-time processing without compromising accuracy. Extensive evaluation demonstrates robustness and high precision under challenging conditions including large motion, flash effects, flickering, low contrast, and noise. The framework delivers reliable, efficient foundational analysis to support higher-level video understanding tasks.
Existing research on video event detection lacks a unified large-scale dataset and standardized evaluation protocols, hindering fair method comparison and reproducibility. To address this gap, this work proposes the first integrated, three-pronged development framework encompassing dataset construction, performance evaluation, and deployment scenarios. By introducing structured data design, a standardized metric system, and diverse application-oriented modeling, the framework establishes a generalizable paradigm for the field. This approach substantially enhances the fairness of algorithmic comparisons, improves research reproducibility, and supports systematic methodological advancement in video event detection.
研究通过分析V-JEPA 2和VideoMAE-v2模型,探讨了视频基础模型中时空表示的编码内容、出现位置及几何组织方式,并使用轻量级探针来发现三种时间属性。
To address the inflexibility, frequent retraining requirements, and difficulty in modeling continuous appearance-motion evolution in video anomaly detection (VAD), this paper proposes a Configurable Spatio-Temporal Hierarchical Architecture (STHA). STHA features a three-level scalable design—flow-level, stack-level, and block-level—enabling on-demand adjustment of detection granularity. It introduces a dual-stream residual stacking mechanism to jointly model normal patterns in RGB spatial frames and optical flow temporal sequences. Furthermore, it integrates multi-capacity anomaly blocks, cross-layer/cross-stack residual connections, and hierarchical normality learning. Evaluated on three mainstream benchmarks—UCSDped2, ShanghaiTech, and CUHK Avenue—STHA achieves state-of-the-art performance. Additionally, experiments on a newly constructed toy dataset demonstrate its adaptive capability to balance diverse detection requirements effectively.
This work addresses the limitation of existing weakly supervised video anomaly detection methods, which primarily focus on temporal localization while lacking precise spatial awareness, thereby hindering interpretability in real-world applications. We propose a patch-based spatiotemporal anomaly localization framework that jointly models the temporal and spatial locations of anomalies using only video-level labels. Leveraging multiple instance learning, our approach infers region-level anomaly scores from grid-level patch features and introduces a novel neighborhood-aware Top-k spatiotemporal selection strategy to generate fine-grained spatial anomaly maps without requiring bounding box supervision. Additionally, we provide frame-level bounding box annotations for two widely used datasets. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches across multiple benchmarks, achieving substantial improvements in spatiotemporal localization accuracy. The code, models, and new annotations are publicly released.
This work addresses the ambiguity regarding whether existing single-stage video object detectors genuinely leverage temporal context, as standard evaluation metrics often fail to reveal their actual reliance on temporal information. To this end, we propose TemporalLens, a diagnostic framework that quantifies a model’s temporal dependency through controlled perturbations—including temporal shuffling, structured occlusion, and redundancy injection. Furthermore, we design YOLO-3D based on YOLOv8, explicitly preserving the temporal dimension within the backbone to enhance genuine temporal reasoning. Experiments demonstrate that TemporalLens effectively distinguishes between stacked 2D models and true temporal architectures, while YOLO-3D achieves an average mAP@50 improvement of 3.7 percentage points with 32-frame inputs, underscoring the critical role of temporal depth in performance gains.
该研究解决了视频中局部运动表示问题,通过全视频处理和基于空间掩码的区域查询方法,生成保留全局上下文的局部动态嵌入。
本文提出AVR,利用预训练模型修复视频中的异常区域,仅在无证据可复制时生成内容,通过实验验证其有效性。
This work addresses the challenge of distinguishing small targets from single-frame backgrounds in infrared image sequences under extremely low signal-to-noise ratios. To this end, we propose a non-interactive segmentation method that integrates temporal motion modeling with the Segment Anything Model (SAM). By characterizing both global motion patterns and local motion deviations, our approach extracts motion discrepancy features to enhance potential target regions and, for the first time, explicitly generates temporally emergent prompts to guide SAM segmentation. This framework effectively combines large-scale semantic pretraining with task-specific temporal cues, significantly improving detection and segmentation performance for small targets in complex dynamic scenes and overcoming the performance limitations of conventional methods under ultra-low signal-to-noise conditions.