PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the coupled challenges of unreliable semantic extraction and high historical transmission overhead in training-free online video anomaly detection by proposing a dual-module framework. The first component, a proposition-aware reasoning module, leverages frozen multimodal large models to extract structured evidence while reusing causal visual prefixes. The second, a streaming evidence upgrade module, maps evidence into compact states to preserve temporal continuity. A core innovation lies in establishing a "current-first" principle that decouples semantic acquisition from score evolution, reducing redundant computation through selective execution and propagating only bounded states to bridge residual gaps. Extensive experiments across four benchmarks demonstrate strong performance, significantly reducing expert computation costs while maintaining efficient sparse state propagation.
📝 Abstract
Training-free online video anomaly detection (VAD) with frozen multimodal language models faces two coupled challenges: extracting reliable current-window semantics under causal and computational constraints, and maintaining temporal continuity without repeatedly transmitting high-dimensional history. Encoding history through text can compress visual evidence and introduce semantic bias, whereas retaining visual history expands multimodal context. We introduce PARSEE-VAD, a two-module framework that separates semantic evidence acquisition from score-state evolution. Proposition-Aware Reasoning (PAR) extracts structured propositional evidence from the current causal window and conditionally activates more specific queries when coarse evidence warrants further refinement. By sharing a reusable causal visual prefix across queries, PAR reduces redundant computation through selective execution. Streaming Evidence Escalation (SEE) maps the acquired proposition evidence into a compact score-domain event state through current evidence escalation, then propagates only the resulting bounded state across decisions to support temporal continuity. Experiments on four benchmarks demonstrate strong training-free online performance while selective routing reduces specialist computation and score-state propagation remains sparse. These results support a current-first principle for streaming multimodal inference: resolve present semantics first, then use compact historical state only to repair residual continuity gaps.
Problem

Research questions and friction points this paper is trying to address.

video anomaly detection
training-free online inference
multimodal language models
temporal continuity
causal constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free Video Anomaly Detection
Proposition-Aware Reasoning
Streaming Evidence Escalation
Multimodal Language Models
Causal Visual Prefix
🔎 Similar Papers
No similar papers found.