Detecting Defects that Matter: An Application-Driven Benchmark for Anomaly Detection in Manufacturing and Retail Logistics (VAND 4.0 Challenge)

πŸ“… 2026-10-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the saturation and limited real-world applicability of existing anomaly detection benchmarks, which inadequately assess defect detection performance and computational efficiency in industrial and retail settings. To this end, we construct an application-driven benchmark featuring hidden test sets for these domains, propose a novel efficiency metric integrating performance with power consumption, and release Kaputt-Rare, a low-prevalence retail dataset. Comprehensive evaluations encompassing pixel-level segmentation and object detection are conducted using DINOv3, vision-language models (VLMs), and specialized supervised detectors. Results reveal that the best-performing method achieves only 57% SegF1 on industrial segmentation, while VLMs lag behind specialized models by approximately 28 AP in retail detection. These findings underscore the persistent challenges associated with efficiency optimization and rare defect detection in practical applications.
πŸ“ Abstract
Existing Anomaly Detection benchmarks are saturated and often unrealistic. As part of the VAND 4.0 Challenge, we introduce a hidden-test, application-driven benchmark across two deployment-critical domains: industrial manufacturing and retail logistics. In the Industrial Track (MVTec AD 2), the results reveal that unsupervised anomaly segmentation remains challenging: the best regular-setting method achieves only ~57\% pixel-level $SegF_1$, indicating substantial room for improvement. Zero-shot approaches trail by ~15 $SegF_1$ points, confirming that task-specific training on normal data remains essential for precise defect localization. Robustness to distribution shifts remains a key open challenge and DINOv3-backbones clearly dominate this track. In the Retail Track (Kaputt 2), the results reveal that (1) supervised defect detection is approaching saturation for common defect types; (2) the best off-the-shelf VLM approach trails specialized models by ~28 AP, confirming that currently VLMs cannot replace fine-tuned detectors, (3) reference images did not prove helpful for top-performing approaches. Performance collapses on rare defects (spillage ~53 AP, missing units ~27 AP), where the supervised ceiling is bounded by data availability. To drive future progress in this domain, we provide a new low-prevalence retail AD dataset (Kaputt-Rare). Across both tracks, computational efficiency is assessed as a first-class metric combining performance, throughput, memory, and power consumption. We introduce a novel metric for measuring efficiency and reveal that that top-performing methods rely on heavy architectures while efficiency is largely neglected. Overall, we conclude that the community needs (a) more efficiency-aware method development, and (b) true anomaly detection approaches for rare defects and shifting conditions. https://sites.google.com/view/vand4-cvpr2026/challenge
Problem

Research questions and friction points this paper is trying to address.

Anomaly Detection
Distribution Shift
Rare Defects
Computational Efficiency
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Anomaly Detection
Application-Driven Benchmark
Computational Efficiency
Low-Prevalence Dataset
Distribution Shift
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
L
Lars Heckler-Kram
MVTec Software GmbH
D
Dorian Henning
Amazon
A
Ashwin Vaidya
Intel
J
Jan-Hendrik Neudeck
MVTec Software GmbH
U
Ulla Scheler
MVTec Software GmbH
Anton Milan
Anton Milan
Amazon
Computer VisionRobotics
Samet Akcay
Samet Akcay
AI Research Engineer at Intel
Computer VisionMachine LearningAnomaly Detection
P
Paula Ramos
NVIDIA
Sebastian HΓΆfer
Sebastian HΓΆfer
Manager Applied Science, Amazon Fulfillment Technologies & Robotics
Computer VisionMachine LearningRobotics