Decision Readouts for Text-Mediated Video Anomaly Detection: An Exploratory Evaluation of Jev and Qwen

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the impact of decision readout mechanisms on video anomaly detection performance when textual evidence is held fixed. Conducting systematic experiments on the UCF-Crime and XD-Violence datasets, we compare multiple readout strategies between Jev-typed decisions and Qwen-generated probabilities. Specifically, we quantitatively analyze differences in ranking capability and probability quality across distinct interfaces under sparse positive sample conditions, and introduce a local ordinal likelihood expectation technique to optimize the decision process. Experimental results demonstrate that the Jev Noul strategy achieves 75.99% AP on the XD dataset, significantly outperforming the Qwen baseline, although it offers no advantage on UCF-Crime. This work reveals the critical role of readout interface design in weakly supervised anomaly detection.
📝 Abstract
How much does the decision readout matter when video-derived textual evidence is held fixed? We evaluate Jev typed decisions and three Qwen readouts on a sparse development sample of 40 videos and 400 target anchors from UCF-Crime and XD-Violence, each presented as a summary and ordered captions. Each dataset contributes 20 source groups and 200 anchors, including only 10 and 37 positives, respectively. The original five-backend pilot requested 4,000 predictions; Jev Choice returned 776 valid responses out of 800 under the study's strict numerical policy, blocking its full-coverage quality comparison. On XD captions, Jev Noul achieved 75.99% average precision versus 48.47% for Qwen generated probability and 57.81% for the stronger local ordinal-likelihood expectation. The latter paired difference was 18.18 percentage points (95% source-group bootstrap interval 5.53-31.50). UCF did not show a corresponding advantage: caption ROC-AUC was 52.26% for Noul and 65.95% for ordinal likelihood. Both probability readouts had higher, hence worse, UCF Brier scores than the evaluation-prevalence reference of 0.0475. We additionally audit historical LAVAD scores at exactly matched anchors and distinguish response structure from numerical consistency. A binary-likelihood control is missing. These exploratory offline results characterize ranking, probability quality and interface failures; they establish neither a causal typed-interface benefit nor general superiority, calibration or end-to-end acceleration.
Problem

Research questions and friction points this paper is trying to address.

Video Anomaly Detection
Decision Readout
Text-Mediated
Large Language Models
Probability Calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Anomaly Detection
Decision Readouts
Text-Mediated Evaluation
Large Language Models
Probability Calibration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xukui Qin
Independent researchers
Y
Youting Wang
Independent researchers
X
Xinjie He
Independent researchers
Ziyang Luo
Ziyang Luo
Salesforce AI Research
AgentsLLMsMultimodal
R
Runxiong Wu
Independent researchers
Y
Yan-Syuan Chen
Independent researchers
Z
Zhongyao Chu
Independent researchers