Representation--Behavior Alignment for Explainable Weakly-Supervised Video Anomaly Detection

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the alignment discrepancy between discriminative information encoded in hidden states and final decisions when applying multimodal large language models to video anomaly detection. By decomposing this misalignment into capacity and directional components, we identify directional bias as the dominant factor and attention layers as the critical bottleneck. Accordingly, we propose a parameter-efficient alignment method that fine-tunes only the attention layers using video-level labels. Experiments demonstrate that this approach significantly improves performance across three benchmarks, precisely aligning decision directions with discriminative representations. Furthermore, it enables synchronized generation of anomaly predictions and interpretable explanations through a single forward pass, achieving a unified framework for weakly supervised detection and generation.
📝 Abstract
Multimodal Large Language Models (MLLMs) provide a natural way to make video anomaly detection more explainable. However, their final decisions do not always fully use the discriminative information contained in their hidden states, an issue we refer to as representation--behavior misalignment. We decompose this gap into a capacity component that measures discriminative information never aggregated into the readout position, and a directional component that measures the angular mismatch between the optimal and the native normal--abnormal axis at that position. Across multiple video anomaly detection benchmarks and MLLM backbones the directional component dominates, and residual-stream tracing shows that native-axis separability rises sharply in several mid-to-late attention layers. Because both components are governed by attention rather than MLP updates, we propose Representation--Behavior Alignment (RBA), a parameter-efficient method that adapts those layers using video-level labels alone while updating about 0.012\% of the backbone parameters. Experiments on three benchmarks show that RBA improves native-readout performance and better aligns the model's decision direction with discriminative representations, and it produces anomaly decisions and explanations through a single generative process.
Problem

Research questions and friction points this paper is trying to address.

Video Anomaly Detection
Representation-Behavior Misalignment
Multimodal Large Language Models
Weakly-Supervised Learning
Explainability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Representation-Behavior Alignment
Weakly-Supervised Video Anomaly Detection
Multimodal Large Language Models
Parameter-Efficient Fine-Tuning
Explainable AI
🔎 Similar Papers
No similar papers found.