Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations

📅 2026-08-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of a systematic synthesis and unified evaluation framework for class activation mapping (CAM) methods across diverse model architectures and eras. Surveying 57 foundational studies since 2016, it proposes a cohesive taxonomy grounded in attribution mechanisms, architectural dependencies, and evaluation objectives, encompassing visualization techniques from CNNs to Transformers and foundation models such as CLIP, DINO, and SAM. The study traces the evolution of CAM from single-layer, low-resolution explanations toward multi-layer, probabilistic, token-aware, and foundation-model-aware paradigms. It systematically summarizes key contributions and persistent challenges across method categories and highlights the current fragmentation in evaluation protocols, offering methodological guidance and a benchmark reference for future research on interpretable vision models.
📝 Abstract
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolutional channels, tokens, or patches that support a target class or concept. Since the first CAM formulation in 2016, the field has moved far beyond global-average-pooled CNN classifiers. CAM-style methods now include gradient-based post-hoc explanations, gradient-free score and ablation methods, high-resolution upscaling, weakly supervised localization and segmentation, transformer token attribution, causal and debiasing methods, and foundation-model-era approaches that use CLIP, DINO, SAM, or feature-distribution comparisons. This review synthesizes a strict corpus of 57 method-centered papers published from 2016 onward. The paper develops a taxonomy that separates methods by attribution mechanism, architectural dependence, and evaluation objective. It then reviews gradient-based CAMs, recent and hybrid CAM-style methods, and model-based or architecture-aware methods. Across the corpus, the main trend is clear: the field is shifting from explaining one class score in one low-resolution CNN layer toward comparative, multi-layer, probabilistic, token-aware, and foundation-model-aware explanations. At the same time, evaluation remains fragmented. Faithfulness, localization, robustness, computational cost, and human trust are often measured with different protocols. The review therefore emphasizes not only what each method contributes, but also which gap it leaves open and which later methods attempt to close that gap.
Problem

Research questions and friction points this paper is trying to address.

Class Activation Mapping
Explainable AI
Visual Explanations
Evaluation Metrics
Foundation Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Class Activation Mapping
Explainable AI
Foundation Models
Visual Explanations
Method Taxonomy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
AmirHossein Eshghi
M.Sc. student, Department of Computer Engineering, University of Birjand, Iran
Hamid Saadatfar
Hamid Saadatfar
Associate Professor, Department of Computer Engineering, University of Birjand, Iran
S
Seyyed Ali Hoseini
Assistant professor, Department of Computer Engineering, University of Birjand, Iran
A
AmirMohsen Eshghi
B.Sc. student, Department of Computer Engineering, University of Birjand, Iran
S
Siavash Arjomand Bigdeli
Associate Professor, Department of Applied Mathematics and Computer Science Visual Computing, Technical University of Denmark, Denmark