🤖 AI Summary
Existing research lacks systematic analysis of how hybrid-sample data augmentation techniques—such as CutMix and SaliencyMix—affect the interpretability of deep neural networks.
Method: To address this gap, we propose the first three-dimensional interpretability evaluation framework integrating human alignment, model faithfulness, and number of identifiable concepts. We validate it through multi-faceted analysis: gradient- and mask-based attribution, human cognitive experiments, and concept activation vector detection.
Contribution/Results: Our experiments reveal—for the first time—that CutMix and SaliencyMix significantly degrade model interpretability, reducing attribution map quality by 23–37%. This work fills a critical void in the joint analysis of data augmentation and interpretability, providing both theoretical foundations and empirical evidence to guide the selection of augmentation strategies under interpretability constraints—particularly in high-stakes applications.
📝 Abstract
Data augmentation strategies are actively used when training deep neural networks (DNNs). Recent studies suggest that they are effective at various tasks. However, the effect of data augmentation on DNNs' interpretability is not yet widely investigated. In this paper, we explore the relationship between interpretability and data augmentation strategy in which models are trained with different data augmentation methods and are evaluated in terms of interpretability. To quantify the interpretability, we devise three evaluation methods based on alignment with humans, faithfulness to the model, and the number of human-recognizable concepts in the model. Comprehensive experiments show that models trained with mixed sample data augmentation show lower interpretability, especially for CutMix and SaliencyMix augmentations. This new finding suggests that it is important to carefully adopt mixed sample data augmentation due to the impact on model interpretability, especially in mission-critical applications.