Faithful Faithfulness Evaluations: Challenges & Pitfalls Learned from a Breast MRI Case Study

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过乳腺MRI案例探讨了基于显著性图的深度学习解释方法可能误导临床医生的问题,评估了多种显著性方法,并指出当前方法存在的挑战及需要更稳健、标准化评估框架的需求。
📝 Abstract
Saliency maps are widely used to explain deep learning predictions in medical imaging, yet visually plausible explanations do not necessarily reflect a model's true decision process and may therefore mislead clinicians. We investigate this problem using a Vision Transformer-based breast MRI classifier trained on the ODELIA Breast MRI Challenge dataset and evaluate multiple saliency methods, including Last-layer Attention, Attention Rollout, Grad-SAM, Gradient Attention Rollout, GMAR, Grad-CAM, and HiResCAM. Our study highlights two often-overlooked challenges in perturbation-based faithfulness evaluation. First, method rankings depend strongly on the perturbation strategy, varying across intensity-based perturbations and transformer-based attention masking. Second, benchmarking saliency methods requires distinguishing between class-specific and class-agnostic explanations. To enable fair comparisons, we introduce non-class-specific variants of gradient-based methods and evaluate both settings separately. Across protocols, Grad-CAM and Gradient Attention Rollout consistently emerged as the strongest class-specific methods, although their relative ranking depended on the evaluation design. These findings expose important limitations of current saliency-based explainability approaches and highlight the need for more robust and standardized evaluation frameworks for trustworthy clinical AI systems.
Problem

Research questions and friction points this paper is trying to address.

saliency maps
medical imaging
decision process
clinicians
faithfulness evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Perturbation Strategy
Class-Specific Explanations
Non-Class-Specific Variants
P
Peachapong Poolpol
Fraunhofer Institute for Digital Medicine MEVIS, Bremen, Germany; Deggendorf Institute of Technology, Deggendorf, Germany
H
Henrik H. J. Detjen
Fraunhofer Institute for Digital Medicine MEVIS, Bremen, Germany
E
Eike Petersen
Fraunhofer Institute for Digital Medicine MEVIS, Bremen, Germany