Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing video anomaly detection methods, which typically rely on task-specific training data, while zero-shot approaches lack temporal continuity and structured reasoning capabilities. To overcome these challenges, this work proposes Cog-VADU, a training-free framework that reformulates anomaly detection as a sequential cognitive reasoning task. Specifically, it introduces the Chain-of-Anomaly-Detection-Thought Prompting (CoADTP) strategy, which maintains implicit temporal memory through recursive reasoning, and incorporates cross-modal re-ranking to enhance semantic consistency. Built upon large vision-language models, the proposed method is model-agnostic and achieves state-of-the-art zero-shot performance across multiple benchmarks. Ultimately, Cog-VADU enables interpretable and strongly generalizable precise anomaly localization and understanding in open-world scenarios.
📝 Abstract
Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but often lack temporal continuity and structured reasoning. We propose Cog-VADU, a fully training-free framework that reformulates VAD as a sequential cognitive reasoning task. Cog-VADU introduces Chain-of- Anomaly Detection Thought Prompting (CoADTP), which unrolls an LVLM into a recurrent reasoning chain across video segments. By propagating structured rationales over time, the model maintains implicit temporal memory, enabling robust discrimination between com- plex anomalies and high-motion normal activities. To improve reliability, we further design a cross-modal re-ranking stage that aligns textual rationales with visual embeddings, enforcing semantic consistency and temporal coherence for refined and stable predictions. Extensive experiments on multiple public VAD benchmarks demonstrate that Cog-VADU achieves competitive zero-shot performance. Moreover, cross-model evaluations show that CoADTP consistently enhances reasoning-based anomaly detection in a model-agnostic manner, pro- viding interpretable and generalizable anomaly understanding for real-world applications.
Problem

Research questions and friction points this paper is trying to address.

Video Anomaly Detection
Zero-shot Generalization
Open-set Scenarios
Temporal Continuity
Structured Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free
Cognitive Reasoning
Video Anomaly Detection
Chain-of-Thought Prompting
Cross-Modal Re-ranking
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mohd Ubaid Wani
Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey, UK
S
Sara Atito
Surrey Institute for People-Centred AI (PAI), University of Surrey, UK
Josef Kittler
Josef Kittler
University of Surrey
engineering
Muhammad Awais
Muhammad Awais
CVSSP, University of Surrey
Self-Supervised LearningMulti-Modal LearningDeep LearningMedical image analysis