video roi localization

Designs and implements algorithms that automatically identify and localize spatial regions of interest in video sequences by analyzing temporal changes (e.g., frame-to-frame or pixel variance) to detect activity or salient areas. Builds modules that output temporally consistent ROI masks or bounding boxes and perform cropping/masking to reduce background and camera effects.

videoroilocalization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Detecting Contextual Anomalies by Discovering Consistent Spatial Regions

Jan 14, 2025
ZY
Zhengye Yang
🏛️ Rensselaer Polytechnic Institute

Addressing the challenges of modeling spatial context, reliance on pre-trained models, and poor interpretability in street-scene video anomaly detection, this paper proposes an unsupervised spatial consistency modeling framework. First, Gaussian Mixture Modeling (GMM) is applied to high-resolution feature maps for unsupervised clustering, jointly discovering object-level spatial attributes and spatially consistent regions. Subsequently, an inter-object spatial relation graph is constructed to generate pixel-level normality heatmaps. The method requires neither pre-trained segmentation models nor human annotations. Evaluated on the Street Scene dataset, it achieves state-of-the-art performance while reducing parameter count by one to two orders of magnitude. Moreover, it produces high-resolution, semantically interpretable anomaly localization maps—enhancing model transparency and enabling efficient deployment.

Anomaly detectionUnsupervised learningVideo analysis

Video foundational analysis requires efficient shot boundary detection, sampling pattern identification, and dynamic keyframe extraction. This paper proposes the first unified framework that simultaneously addresses hard-cut and short gradual shot boundary detection, discriminates scan-based sampling modes (progressive, interlaced, and film stretch), and performs motion-adaptive keyframe selection. Our method innovatively integrates motion field estimation with normalized cross-correlation (NCC) features, jointly models inter-frame and intra-frame multi-scale features, and introduces a sparse selective computation strategy—achieving 4× real-time processing without compromising accuracy. Extensive evaluation demonstrates robustness and high precision under challenging conditions including large motion, flash effects, flickering, low contrast, and noise. The framework delivers reliable, efficient foundational analysis to support higher-level video understanding tasks.

Analyzes sampling structure in videosDetects video shot boundaries efficientlyIdentifies dynamic keyframes faster than real-time

Existing research on video event detection lacks a unified large-scale dataset and standardized evaluation protocols, hindering fair method comparison and reproducibility. To address this gap, this work proposes the first integrated, three-pronged development framework encompassing dataset construction, performance evaluation, and deployment scenarios. By introducing structured data design, a standardized metric system, and diverse application-oriented modeling, the framework establishes a generalizable paradigm for the field. This approach substantially enhances the fairness of algorithmic comparisons, improves research reproducibility, and supports systematic methodological advancement in video event detection.

datasetevent detectionmethod comparison

Configurable Spatial-Temporal Hierarchical Analysis for Flexible Video Anomaly Detection

May 12, 2023
KC
Kai Cheng
🏛️ Fudan University | East China University of Science and Technology | Shanghai University of Traditional Chinese Medicine

To address the inflexibility, frequent retraining requirements, and difficulty in modeling continuous appearance-motion evolution in video anomaly detection (VAD), this paper proposes a Configurable Spatio-Temporal Hierarchical Architecture (STHA). STHA features a three-level scalable design—flow-level, stack-level, and block-level—enabling on-demand adjustment of detection granularity. It introduces a dual-stream residual stacking mechanism to jointly model normal patterns in RGB spatial frames and optical flow temporal sequences. Furthermore, it integrates multi-capacity anomaly blocks, cross-layer/cross-stack residual connections, and hierarchical normality learning. Evaluated on three mainstream benchmarks—UCSDped2, ShanghaiTech, and CUHK Avenue—STHA achieves state-of-the-art performance. Additionally, experiments on a newly constructed toy dataset demonstrate its adaptive capability to balance diverse detection requirements effectively.

Creates a compatible dataset to evaluate detection performance under varying demandsDesigns configurable video anomaly detection to avoid retraining for demand changesDevelops a module for multi-scale spatial-temporal coherence to improve accuracy

Latest Papers

What's happening recently
View more

This work addresses the limitation of existing weakly supervised video anomaly detection methods, which primarily focus on temporal localization while lacking precise spatial awareness, thereby hindering interpretability in real-world applications. We propose a patch-based spatiotemporal anomaly localization framework that jointly models the temporal and spatial locations of anomalies using only video-level labels. Leveraging multiple instance learning, our approach infers region-level anomaly scores from grid-level patch features and introduces a novel neighborhood-aware Top-k spatiotemporal selection strategy to generate fine-grained spatial anomaly maps without requiring bounding box supervision. Additionally, we provide frame-level bounding box annotations for two widely used datasets. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches across multiple benchmarks, achieving substantial improvements in spatiotemporal localization accuracy. The code, models, and new annotations are publicly released.

Anomaly DetectionSpatial LocalizationSpatiotemporal Localization

This work addresses the ambiguity regarding whether existing single-stage video object detectors genuinely leverage temporal context, as standard evaluation metrics often fail to reveal their actual reliance on temporal information. To this end, we propose TemporalLens, a diagnostic framework that quantifies a model’s temporal dependency through controlled perturbations—including temporal shuffling, structured occlusion, and redundancy injection. Furthermore, we design YOLO-3D based on YOLOv8, explicitly preserving the temporal dimension within the backbone to enhance genuine temporal reasoning. Experiments demonstrate that TemporalLens effectively distinguishes between stacked 2D models and true temporal architectures, while YOLO-3D achieves an average mAP@50 improvement of 3.7 percentage points with 32-frame inputs, underscoring the critical role of temporal depth in performance gains.

model diagnosticssingle-stage detectorstemporal context

This work addresses the challenge of distinguishing small targets from single-frame backgrounds in infrared image sequences under extremely low signal-to-noise ratios. To this end, we propose a non-interactive segmentation method that integrates temporal motion modeling with the Segment Anything Model (SAM). By characterizing both global motion patterns and local motion deviations, our approach extracts motion discrepancy features to enhance potential target regions and, for the first time, explicitly generates temporally emergent prompts to guide SAM segmentation. This framework effectively combines large-scale semantic pretraining with task-specific temporal cues, significantly improving detection and segmentation performance for small targets in complex dynamic scenes and overcoming the performance limitations of conventional methods under ultra-low signal-to-noise conditions.

infrared small target detectionlow signal-to-noise ratiomultiframe segmentation

Hot Scholars

MN

Mahdi Nikdast

Electrical and Computer Engineering, Colorado State University
Silicon PhotonicsIntegrated PhotonicsTest and VerificationHigh-Performance Computing
SA

Shaahin Angizi

Assistant Professor at New Jersey Institute of Technology
In-Memory ComputingIn-Sensor ComputingMemory SecurityAI
JH

Jana Hutter

UKER/FAU Erlangen // King's College London
Magnetic Resonance ImagingPerinatal ImagingQuantitative Imaging
GD

Gourav Datta

Assistant Professor, Case Western Reserve University
CZ

Chengwei Zhou

Zhejiang University
Array Signal ProcessingDirection-of-Arrival EstimationRobust Adaptive Beamforming