π€ AI Summary
This study addresses the saturation and limited real-world applicability of existing anomaly detection benchmarks, which inadequately assess defect detection performance and computational efficiency in industrial and retail settings. To this end, we construct an application-driven benchmark featuring hidden test sets for these domains, propose a novel efficiency metric integrating performance with power consumption, and release Kaputt-Rare, a low-prevalence retail dataset. Comprehensive evaluations encompassing pixel-level segmentation and object detection are conducted using DINOv3, vision-language models (VLMs), and specialized supervised detectors. Results reveal that the best-performing method achieves only 57% SegF1 on industrial segmentation, while VLMs lag behind specialized models by approximately 28 AP in retail detection. These findings underscore the persistent challenges associated with efficiency optimization and rare defect detection in practical applications.
π Abstract
Existing Anomaly Detection benchmarks are saturated and often unrealistic. As part of the VAND 4.0 Challenge, we introduce a hidden-test, application-driven benchmark across two deployment-critical domains: industrial manufacturing and retail logistics. In the Industrial Track (MVTec AD 2), the results reveal that unsupervised anomaly segmentation remains challenging: the best regular-setting method achieves only ~57\% pixel-level $SegF_1$, indicating substantial room for improvement. Zero-shot approaches trail by ~15 $SegF_1$ points, confirming that task-specific training on normal data remains essential for precise defect localization. Robustness to distribution shifts remains a key open challenge and DINOv3-backbones clearly dominate this track. In the Retail Track (Kaputt 2), the results reveal that (1) supervised defect detection is approaching saturation for common defect types; (2) the best off-the-shelf VLM approach trails specialized models by ~28 AP, confirming that currently VLMs cannot replace fine-tuned detectors, (3) reference images did not prove helpful for top-performing approaches. Performance collapses on rare defects (spillage ~53 AP, missing units ~27 AP), where the supervised ceiling is bounded by data availability. To drive future progress in this domain, we provide a new low-prevalence retail AD dataset (Kaputt-Rare). Across both tracks, computational efficiency is assessed as a first-class metric combining performance, throughput, memory, and power consumption. We introduce a novel metric for measuring efficiency and reveal that that top-performing methods rely on heavy architectures while efficiency is largely neglected. Overall, we conclude that the community needs (a) more efficiency-aware method development, and (b) true anomaly detection approaches for rare defects and shifting conditions. https://sites.google.com/view/vand4-cvpr2026/challenge