Spatial Attention Supervision for Defect Localization: Exploiting Ground-Truth Masks as Training Signal in Diffusion-Augmented Defect Detection

📅 2026-09-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过将真实缺陷掩码作为训练信号,并结合扩散模型增强,改进了工业检测中缺陷定位的精度。
📝 Abstract
Ground-truth defect masks in industrial inspection datasets are typically reserved for evaluation. This paper repurposes them as spatial supervision signals during training of classification networks, teaching a model not just what to predict but where to look. The method adds an activation-based attention alignment loss that steers convolutional feature maps toward defect regions, in a mixed-supervision formulation that also accommodates samples without masks, such as diffusion-generated images. Combined with DDPM augmentation, synthetic images contribute quantity while masks contribute spatial precision. We evaluate 85 models (four CNN backbones under a 2x2 data/training factorial over five seeds, plus a Swin-V2-T transformer baseline) on the MVTec-AD bottle benchmark, with localization measured on held-out defect images excluded from classifier gradient updates. Main findings: (1) attention-guided training improves activation-based localization (Pixel-AUROC) by +18.0% for EfficientNetB0 with augmentation (p=0.005, Cohen's d=2.6) and +18.7% for ResNet50 (p=0.008), significant in four of eight CNN settings (uncorrected for multiple comparisons) with no significant change in classification; (2) for EfficientNetB0 a data x training-mode interaction is significant (p=0.002), consistent with a super-additive effect (+13.6% combined vs +1.6% summed individual effects); (3) architectures with weaker spatial representations benefit most, whereas ConvNeXt-T shows no effect, apparently because its depthwise-convolution activations yield spatially uninformative channel-mean maps; (4) unsupervised PatchCore remains the strongest localizer (Pixel-AUROC=0.983), contextualizing the supervised gains. These results show that existing evaluation masks can act as practical training signals that measurably and reproducibly improve where defect classifiers attend.
Problem

Research questions and friction points this paper is trying to address.

Defect Localization
Ground-Truth Masks
Spatial Supervision
Attention Alignment
Industrial Inspection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Attention Supervision
Ground-Truth Masks
Activation-based Attention Alignment Loss
Mixed-Supervision
DDPM Augmentation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Sajjad Rezvani Boroujeni
Sajjad Rezvani Boroujeni
Bowling Green State University
Data Science Machine Learning
M
Muskan Saraf
Data Science Department, Actual Reality Technologies, OH, USA.
G
Gnana Tulasi Makineni
Data Science Department, Actual Reality Technologies, OH, USA.
T
Tom Bush
Data Science Department, Actual Reality Technologies, OH, USA.
H
Hossein Abedi
Data Science Department, Actual Reality Technologies, OH, USA.