Comparing Post-Hoc Explainable AI Methods for Interpreting Black-Box EEG Models in Depression Detection

📅 2026-05-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Deep learning demonstrates superior performance in EEG-based depression detection, yet its black-box nature hinders clinical adoption. This study presents the first systematic comparison of three families of post-hoc explainability methods—Shapley-value-based (DeepSHAP), gradient-based (Integrated Gradients, GradCAM), and perturbation-based (Occlusion, Permutation Feature Importance)—evaluating their attribution consistency and neurophysiological plausibility on EEG time-series data using the InceptionTime model and subject-wise stratified cross-validation. Results reveal high agreement between gradient- and perturbation-based methods, both highlighting right-hemisphere frontal, temporal, and posterior regions. Although DeepSHAP broadly aligns with prior knowledge of major depressive disorder (MDD), it exhibits markedly distinct attribution patterns, underscoring the critical influence of explanation method choice on interpretability outcomes.
📝 Abstract
Recent advances in deep learning have enabled increasingly accurate electroencephalography (EEG)-based classification of Major Depressive Disorder (MDD), but the decision-making processes of high-capacity models remain difficult to interpret. This study investigates multiple post-hoc explainability methods applied to an InceptionTime architecture trained for EEG-based MDD detection. The analysis includes Shapley-based, gradient-based, and perturbation-based attribution approaches: DeepSHAP, Integrated Gradients, GradCAM, Occlusion, and Permutation Feature Importance. Explainability analysis was performed within a subject-level stratified 5-fold cross-validation framework using global attribution aggregation across EEG segments and subjects. The evaluated methods revealed partially convergent attribution patterns, with recurring emphasis on frontal, temporal, and posterior EEG regions, particularly in the right hemisphere. Quantitative comparison demonstrated substantial agreement between gradient- and perturbation-based approaches, while DeepSHAP produced comparatively distinct attribution distributions. At the same time, variability between explainability methods highlighted the influence of methodological assumptions on the resulting explanations. Overall, the results suggest that different post-hoc explainability approaches capture partially overlapping relevance structures in EEG-based deep learning models for depression detection. Although the observed attribution patterns are broadly consistent with several previous EEG studies of MDD, the analysis should be interpreted as exploratory rather than evidence of definitive neurophysiological biomarkers or clinical applicability. The study highlights both the usefulness and limitations of post-hoc explainability for interpreting black-box EEG classifiers in psychiatric applications.
Problem

Research questions and friction points this paper is trying to address.

Explainable AI
EEG
Depression Detection
Black-Box Models
Post-Hoc Explanation
Innovation

Methods, ideas, or system contributions that make the work stand out.

post-hoc explainability
EEG-based depression detection
InceptionTime
attribution methods
cross-validation framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Antonia Šarčević
University of Zagreb Faculty of Electrical Engineering and Computing, Croatia
N
Nikolina Frid
University of Zagreb Faculty of Electrical Engineering and Computing, Croatia