🤖 AI Summary
This study addresses the limitations of existing post-hoc explanation methods—such as LIME and SHAP—in clinical natural language processing, particularly their inability to generate semantically coherent and clinically meaningful interpretations when applied to long, unstructured medical narratives. Focusing on the task of predicting hospital length of stay, the work presents the first systematic evaluation of these methods’ robustness and reliability on real-world clinical texts through adversarial input perturbations and attribution stability analysis. The findings reveal that current approaches frequently overemphasize irrelevant terms, produce unstable attributions, and yield high-confidence predictions even for nonsensical inputs. These shortcomings expose critical flaws for clinical deployment and underscore the necessity for explanations that simultaneously uphold clinical relevance, semantic consistency, and resilience to linguistic noise.
📝 Abstract
Explaining the predictions of neural models in clinical NLP remains a significant challenge, especially for complex tasks involving long, unstructured medical texts. While post-hoc methods like LIME and SHAP are widely used, they often fall short when applied to clinical narratives. In this paper, we identify core limitations of token-level and perturbation-based explanation techniques through targeted demonstra- tions on a hospital length-of-stay prediction task. Our findings reveal issues such as overemphasis on non-informative tokens, instability in at- tributions, and high-confidence predictions for incoherent input variants. These results underscore the need for explanation strategies that are clin- ically meaningful, semantically grounded, and robust to linguistic noise.