🤖 AI Summary
This study addresses the coupled challenges of unreliable semantic extraction and high historical transmission overhead in training-free online video anomaly detection by proposing a dual-module framework. The first component, a proposition-aware reasoning module, leverages frozen multimodal large models to extract structured evidence while reusing causal visual prefixes. The second, a streaming evidence upgrade module, maps evidence into compact states to preserve temporal continuity. A core innovation lies in establishing a "current-first" principle that decouples semantic acquisition from score evolution, reducing redundant computation through selective execution and propagating only bounded states to bridge residual gaps. Extensive experiments across four benchmarks demonstrate strong performance, significantly reducing expert computation costs while maintaining efficient sparse state propagation.
📝 Abstract
Training-free online video anomaly detection (VAD) with frozen multimodal language models faces two coupled challenges: extracting reliable current-window semantics under causal and computational constraints, and maintaining temporal continuity without repeatedly transmitting high-dimensional history. Encoding history through text can compress visual evidence and introduce semantic bias, whereas retaining visual history expands multimodal context. We introduce PARSEE-VAD, a two-module framework that separates semantic evidence acquisition from score-state evolution. Proposition-Aware Reasoning (PAR) extracts structured propositional evidence from the current causal window and conditionally activates more specific queries when coarse evidence warrants further refinement. By sharing a reusable causal visual prefix across queries, PAR reduces redundant computation through selective execution. Streaming Evidence Escalation (SEE) maps the acquired proposition evidence into a compact score-domain event state through current evidence escalation, then propagates only the resulting bounded state across decisions to support temporal continuity. Experiments on four benchmarks demonstrate strong training-free online performance while selective routing reduces specialist computation and score-state propagation remains sparse. These results support a current-first principle for streaming multimodal inference: resolve present semantics first, then use compact historical state only to repair residual continuity gaps.