Spatial Supervision Without Attribution Optimization: Improving Post-Hoc Class Activation Maps via Box-Guided Evidence Routing

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient spatial localization precision of posterior class activation maps (CAMs) resulting from conventional training. To overcome this limitation, it proposes Box-Guided Evidence Routing (BGER), which leverages bounding box supervision to train a lightweight gating module that routes classification features, optimizing ResNet and DenseNet backbones while maintaining independent Grad-CAM evaluation. Notably, this work is the first to enhance CAM quality by reshaping the backbone network using only inexpensive spatial supervision without directly optimizing attribution maps. Experimental results demonstrate that BGER significantly improves MaxBoxAccV2 from 0.584 to 0.715 on CUB and from 0.757 to 0.832 on Stanford Dogs, while preserving comparable classification accuracy.
📝 Abstract
Post-hoc class activation maps (CAMs) are a standard tool for inspecting the evidence behind an image classifier's predictions, yet nothing in ordinary training encourages these maps to be spatially appropriate. We study whether inexpensive spatial supervision can improve a classifier's own predicted-class Grad-CAM without ever optimizing an attribution map. Box-Guided Evidence Routing (BGER) trains a lightweight gate on the final feature map under box or mask supervision and routes classification through the gated features, while Grad-CAM is computed separately at the pre-gate representation, so the evaluated map never enters the training objective. With a BCE routing loss, BGER raises MaxBoxAccV2 from $0.584$ to $0.715$ on CUB-200-2011 and from $0.757$ to $0.832$ on Stanford Dogs at comparable accuracy. Matched controls attribute most of the ResNet-50 gain to the spatial supervision reshaping the backbone rather than to routing itself: when classification bypasses the gate, most of the improvement remains, and detaching gradients through the gate leaves the ResNet-50 result nearly unchanged. The same detachment preserves most of the gain in two DenseNet-121 chest X-ray settings but removes the apparent gain on Swin-T, and directly supervising the CAM reaches stronger localization at a larger accuracy cost. Overall, spatial supervision can improve separately evaluated post-hoc CAMs, but both the mechanism and the size of the benefit depend on the architecture and the evaluation setting.
Problem

Research questions and friction points this paper is trying to address.

post-hoc class activation maps
spatial supervision
attribution optimization
image classification
localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Box-Guided Evidence Routing
Post-hoc CAMs
Spatial Supervision
Grad-CAM
Attribution Optimization