🤖 AI Summary
Deep learning excels in image analysis but suffers from poor interpretability, hindering its trustworthy deployment in safety-critical applications. This paper systematically surveys four mainstream explainable AI (xAI) paradigms in computer vision: saliency maps, concept bottleneck models, prototype-driven methods, and hybrid approaches—unifying their underlying mechanisms, applicability boundaries, and inherent limitations. We propose a novel, multi-dimensional evaluation framework integrating fidelity, stability, and human agreement to quantitatively compare explanation quality and computational cost across methods. Our key contribution is a task-aware xAI selection guideline—the first structured framework encompassing theoretical foundations, technical pathways, and empirical validation. This work advances model transparency and provides systematic support for deploying xAI in high-stakes visual domains such as medical imaging and surveillance.
📝 Abstract
Deep learning has become the de facto standard and dominant paradigm in image analysis tasks, achieving state-of-the-art performance. However, this approach often results in "black-box" models, whose decision-making processes are difficult to interpret, raising concerns about reliability in critical applications. To address this challenge and provide human a method to understand how AI model process and make decision, the field of xAI has emerged. This paper surveys four representative approaches in xAI for visual perception tasks: (i) Saliency Maps, (ii) Concept Bottleneck Models (CBM), (iii) Prototype-based methods, and (iv) Hybrid approaches. We analyze their underlying mechanisms, strengths and limitations, as well as evaluation metrics, thereby providing a comprehensive overview to guide future research and applications.