Rethinking Global Average Pooling: Your Classifier Is Secretly a Multi-Instance Learner

📅 2026-06-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Standard image classifiers employing global average pooling (GAP) discard spatial information, making it difficult to localize class-discriminative evidence in multi-object scenes. This work reveals, for the first time, that the conventional architecture—comprising GAP followed by a linear classification head—inherently exhibits multi-instance learning (MIL) characteristics, naturally treating an image as a bag of spatial instances. Leveraging this insight, we propose a post-hoc method that requires no model modification and recovers localized class evidence obscured by pooling through predictive grid decomposition. Experiments demonstrate that our approach effectively reconstructs faithful foreground responses using off-the-shelf classifiers and further uncovers that classification failures often stem from the intrinsic limitations of mean aggregation inherent in GAP.
📝 Abstract
Modern image classifiers widely adopt global average pooling (GAP) followed by a linear classification head. This linearity ensures that the image-level logits equal the average of logits obtained by applying the classification head pointwise to the feature grid prior to GAP. Consequently, standard classifiers may inherently retain spatial class evidence that remains recoverable even when the image-level prediction is incorrect. This structure naturally suggests a multiple-instance learning (MIL) interpretation, where an image is viewed as a bag of spatial instances. Within this formulation, we demonstrate that standard classifiers trained with a single label per image can still learn the intended classification task in multi-object scenes. We further exploit this property to decompose image-level logits into a prediction grid, providing a post-hoc diagnostic to extract spatial class evidence that GAP otherwise obscures. Our systematic evaluation reveals that off-the-shelf models consistently recover the ground-truth class within foreground regions. The MIL interpretation further suggests that common classifier failures reflect known limitations of mean aggregation.
Problem

Research questions and friction points this paper is trying to address.

global average pooling
multiple-instance learning
image classification
spatial class evidence
mean aggregation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Global Average Pooling
Multiple Instance Learning
Spatial Class Evidence
Post-hoc Interpretability
Mean Aggregation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Aray Karjauv