Analysis of the Shortcut Learning and Clever Hans Effect in CNN based ECG Image Classification

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether electrocardiogram (ECG) image classification models rely on non-physiological visual artifacts rather than genuine waveform morphology, potentially undermining clinical trustworthiness. The authors construct six controlled image datasets and employ convolutional neural networks alongside Integrated Gradients attribution, occlusion sensitivity analysis, and diverse image perturbations—including cropping, masking, and blurring—to systematically evaluate the presence of shortcut learning and Clever Hans effects. They introduce quantitative metrics such as “shortcut retention score” and “prediction consistency,” complemented by attribution analyses to elucidate model decision rationales. Experimental results demonstrate that certain models maintain high performance even when waveform information is removed or synthetic artifacts are introduced, confirming their dependence on non-clinical cues and highlighting significant interpretability risks in current approaches.
📝 Abstract
Deep learning models for ECG image classification may achieve high accuracy by exploiting non-physiological visual cues instead of ECG waveform morphology. Given the black-box nature of deep learning models, their promise of high predictive performance often remains insufficiently translated into clinical or real-world trust, interpretability, and actionable decision-making. In this study, we examine shortcut learning and Clever Hans effect in a publicly available ECG image dataset using convolutional neural networks. In process we have created six image-derived feature sets (FSs), FS1: raw full ECG images, FS2: cropped waveform-only images, FS3: waveform-masked metadata images, FS4: red-arrow artifact images for the myocardial infarction class, FS5: contrast-enhanced images for the abnormal heartbeat class and FS6: Gaussian-blurred images for the normal class. These controlled representations were used to test whether classification performance persists when waveform information is removed or when artificial class-specific artifacts are introduced. Shortcut retention score, prediction consistency and confidence divergence across Feature-Set Representations have been calculated to assess the transparency about the learning pattern. Along with factual results, average Integrated Gradients and occlusion sensitivity test results are presented to inspect whether model attribution focused on ECG-relevant waveform regions or on non-clinical artifacts. Performance changes across feature sets and attribution patterns were used to identify potential Clever Hans behavior. This study evaluates whether ECG image classifiers learn clinically meaningful morphology or shortcut cues introduced by report layout, metadata, contrast, blur, or artificial markers.
Problem

Research questions and friction points this paper is trying to address.

shortcut learning
Clever Hans effect
ECG image classification
convolutional neural networks
model interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

shortcut learning
Clever Hans effect
ECG image classification
model interpretability
feature-set ablation
🔎 Similar Papers
No similar papers found.
A
Abhay Kumar Pathak
Department of Computer Science, Institute of Science, Banaras Hindu University, Varanasi, India; Department of Computer Science (IDI), Norwegian University of Science and Technology (NTNU), Gjøvik, Norway
M
Mrityunjay Chaubey
University of Petroleum and Energy Studies, Dehradun, India
M
Manjari Gupta
Department of Computer Science, Institute of Science, Banaras Hindu University, Varanasi, India
Deepti Mishra
Deepti Mishra
Department of Computer Science, NTNU-Norwegian University of Science and Technology
Empirical Software EngineeringHuman Robot InteractionEducational Technology