local motion deviation detection

Designs and implements algorithms and systems that detect localized deviations from expected motion patterns in spatiotemporal signals; builds detectors that analyze short-range optical flow, track trajectories, or sensor motion streams to localize, time, and quantify anomalous movement. Evaluates and refines these methods to handle noise, occlusion, varying scales, and temporal dynamics in local motion data.

localmotiondeviationdetection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitation of existing weakly supervised video anomaly detection methods, which primarily focus on temporal localization while lacking precise spatial awareness, thereby hindering interpretability in real-world applications. We propose a patch-based spatiotemporal anomaly localization framework that jointly models the temporal and spatial locations of anomalies using only video-level labels. Leveraging multiple instance learning, our approach infers region-level anomaly scores from grid-level patch features and introduces a novel neighborhood-aware Top-k spatiotemporal selection strategy to generate fine-grained spatial anomaly maps without requiring bounding box supervision. Additionally, we provide frame-level bounding box annotations for two widely used datasets. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches across multiple benchmarks, achieving substantial improvements in spatiotemporal localization accuracy. The code, models, and new annotations are publicly released.

Anomaly DetectionSpatial LocalizationSpatiotemporal Localization

This work addresses the challenge of monitoring robotic task execution when failure modes cannot be exhaustively predefined. We propose a general-purpose, enumeration-free visual anomaly detection method that jointly models camera motion, robot kinematics, and learned optical flow prediction. By fusing observed and predicted optical flow, kinematic constraints, and 3D rigid-body transformation errors, we construct a multi-source anomaly score. A probabilistic U-Net enables uncertainty-aware optical flow prediction, and threshold-based decisioning ensures real-time anomaly detection. Our key contribution is the first unified integration of visual, kinematic, and geometric priors for unsupervised anomaly detection. Evaluated on a book-placing task, our method achieves an AUC of 0.804 and an AP of 0.549—substantially outperforming existing baselines.

Detecting visual anomalies during robot task executionImproving anomaly detection accuracy in roboticsPredicting motion errors for failure identification

Weakly supervised video anomaly detection is often hindered by interference from background clutter and scene-level cues, leading to spatial bias and limited interpretability. To address this, this work proposes the SST-WSVAD framework, which dynamically sparsifies attention to focus on critical spatiotemporal regions and integrates end-to-end coupled spatial and temporal branches for fine-grained anomaly localization. The method innovatively introduces a motion-aware regularization mechanism that operates without external detectors or language prompts, enabling patch-level analysis of spatial bias and facilitating auditable evaluation. Evaluated on UCF-Crime, XD-Violence, and MSAD datasets, the approach achieves performance on par with state-of-the-art methods while offering fine-grained interpretability in both anomaly localization and scene-induced bias.

background biasethical concernsspatial localization

Anomaly Detection in Nonstationary Videos Using Time-Recursive Differencing Network-Based Prediction

Mar 04, 2025
GV
Gargi V. Pillai
🏛️ Indian Institute of Technology

To address the limited anomaly detection performance in non-stationary videos—such as aerial remote sensing sequences—this paper explicitly models temporal non-stationarity for the first time. We propose the Temporal Recursive Differencing Network (TRDN), which embeds differential preprocessing into a deep predictive backbone and jointly employs autoregressive moving-average (ARMA) estimation for dynamic statistical modeling. TRDN further integrates optical flow–based feature extraction with a prediction-error-driven anomaly scoring mechanism. Evaluated on three aerial video datasets and two standard video anomaly detection (VAD) benchmarks, our method achieves state-of-the-art (SOTA) performance, surpassing existing approaches in both Equal Error Rate (EER) and Area Under the Curve (AUC). It significantly enhances robustness to time-varying feature distributions and improves fine-grained anomaly localization accuracy.

Detects anomalies in non-stationary video data.Handles time-varying feature statistics effectively.Improves prediction using time-recursive differencing network.

Configurable Spatial-Temporal Hierarchical Analysis for Flexible Video Anomaly Detection

May 12, 2023
KC
Kai Cheng
🏛️ Fudan University | East China University of Science and Technology | Shanghai University of Traditional Chinese Medicine

To address the inflexibility, frequent retraining requirements, and difficulty in modeling continuous appearance-motion evolution in video anomaly detection (VAD), this paper proposes a Configurable Spatio-Temporal Hierarchical Architecture (STHA). STHA features a three-level scalable design—flow-level, stack-level, and block-level—enabling on-demand adjustment of detection granularity. It introduces a dual-stream residual stacking mechanism to jointly model normal patterns in RGB spatial frames and optical flow temporal sequences. Furthermore, it integrates multi-capacity anomaly blocks, cross-layer/cross-stack residual connections, and hierarchical normality learning. Evaluated on three mainstream benchmarks—UCSDped2, ShanghaiTech, and CUHK Avenue—STHA achieves state-of-the-art performance. Additionally, experiments on a newly constructed toy dataset demonstrate its adaptive capability to balance diverse detection requirements effectively.

Creates a compatible dataset to evaluate detection performance under varying demandsDesigns configurable video anomaly detection to avoid retraining for demand changesDevelops a module for multi-scale spatial-temporal coherence to improve accuracy

Latest Papers

What's happening recently
View more

This work addresses the challenge of efficiently localizing an unknown number of anomalous regions in large-scale spatially dependent data. The authors propose SPLADE, a two-stage method that integrates intelligent sampling with boundary estimation to simultaneously and consistently estimate both the number and boundaries of multiple axis-aligned anomalous patches under general spatial dependence structures—without requiring full spatial grid segmentation. By leveraging a uniform Gaussian approximation and an efficient search strategy, SPLADE substantially improves computational efficiency and localization accuracy. Experimental results demonstrate that SPLADE outperforms existing approaches on both synthetic and real-world video surveillance datasets, achieving faster runtime, higher localization precision, and robustness to strong spatial dependencies.

anomalous patcheslocalizationmultiple anomalies

This work proposes a fully fixed-point, non-iterative, streaming optical flow algorithm tailored for efficient deployment on resource-constrained FPGAs. By partitioning asynchronous event streams into fixed-time windows and representing them as 1-bit spatial occupancy grids, the method evaluates multiple velocity hypotheses in parallel using only integer logic—comprising shift registers, counters, comparators, and LUT-based multipliers—without requiring frame reconstruction, floating-point arithmetic, or division operations. A single-axis prototype was successfully implemented on a Xilinx Artix-7 FPGA, occupying less than 2 kB of memory and achieving 99.5% directional accuracy under event densities of 10–40%. To the best of our knowledge, this is the first demonstration of low-latency, sparse velocity estimation on an FPGA without relying on DSP blocks or dedicated dividers.

event-based visionFPGA implementationhardware efficiency

This work addresses the limitation of existing skeleton-based video anomaly detection methods in simultaneously modeling discrete semantic primitives and fine-grained motion details of human activities, which constrains their multi-level anomaly discrimination capability. To overcome this, the paper introduces a hierarchical motion semantic modeling framework that decomposes skeleton sequences into interpretable semantic primitives and pose details. These components are jointly modeled using a vector-quantized variational autoencoder, an autoregressive Transformer, and conditional normalizing flows, enabling high-precision anomaly detection while preserving privacy. The proposed method achieves state-of-the-art performance with AUC scores of 88.1% on HR-ShanghaiTech and 75.8% on HR-UBnormal, significantly enhancing multi-granularity anomaly recognition.

hierarchical motion semanticsmotion primitivesprivacy-preserving

This work addresses a critical limitation in existing video anomaly detection methods, which rely on frame-level evaluation metrics that fail to capture the ability to detect continuous anomalous events in real-world scenarios, thereby inflating performance estimates. To remedy this, the authors propose an event-centric paradigm: they first analyze the event structure of mainstream datasets and establish the first event-level evaluation benchmark. They then introduce two event localization strategies—a post-processing pipeline based on hierarchical Gaussian smoothing and adaptive binarization, and an end-to-end dual-branch network—incorporating temporal action localization techniques such as tIoU matching and multi-threshold F1 scoring. Experiments reveal that while current state-of-the-art models achieve over 52% frame-level AUC-ROC on datasets like NWPUC, their event-level localization accuracy (at tIoU=0.2) falls below 10%, with an average event F1 score of merely 0.11, underscoring the necessity of shifting to event-level evaluation and highlighting the contributions of this study.

Event-level EvaluationFrame-level MetricsHuman-Centric Surveillance

Hot Scholars

MA

Melkamu Abay Mersha

PhD Candidate, University of Colorado Colorado Springs
AIMachine learningXAINLP
JK

Jugal Kalita

University of Colorado, Colorado Springs
Natural Language ProcessingComputational LinguisticsAnomaly DetectionCybersecurity
AK

Ali K. AlShami

University of Colorado Colorado Springs (UCCS)
Computer Vision/Deep Learning
YS

Yousra Shleibik

Graduate Research Assistant at The University of Denver
Machine LearningComputer VisionHRIRobotics