Explainability for Vision Foundation Models: A Survey

📅 2025-01-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Visual foundation models (e.g., ViT, SAM, diffusion models) suffer from limited interpretability, hindering trust and deployment. Method: This work establishes the first comprehensive XAI research framework tailored to visual foundation models, leveraging bibliometric analysis and topic modeling to systematically categorize and critically assess attribution methods, concept activation, surrogate models, and visualization techniques—explicitly distinguishing their dual roles as *targets of explanation* and *explanation tools*. It identifies the fundamental tension among scalability, fidelity, and generalizability, and proposes standardized evaluation dimensions. Contribution/Results: The study delivers the first XAI research map for visual foundation models, offering a structured taxonomy, critical insights into methodological trade-offs, and foundational guidelines for advancing both theoretical understanding and practical implementation of explainability in this rapidly evolving domain.

Technology Category

Computer Vision: Interpretability, Explainability, and TransparencyHumans and AI: Explainable AI (XAI) for Human UnderstandingMachine Learning: Transparent, Interpretable, Explainable ML

Application Category

User Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
As artificial intelligence systems become increasingly integrated into daily life, the field of explainability has gained significant attention. This trend is particularly driven by the complexity of modern AI models and their decision-making processes. The advent of foundation models, characterized by their extensive generalization capabilities and emergent uses, has further complicated this landscape. Foundation models occupy an ambiguous position in the explainability domain: their complexity makes them inherently challenging to interpret, yet they are increasingly leveraged as tools to construct explainable models. In this survey, we explore the intersection of foundation models and eXplainable AI (XAI) in the vision domain. We begin by compiling a comprehensive corpus of papers that bridge these fields. Next, we categorize these works based on their architectural characteristics. We then discuss the challenges faced by current research in integrating XAI within foundation models. Furthermore, we review common evaluation methodologies for these combined approaches. Finally, we present key observations and insights from our survey, offering directions for future research in this rapidly evolving field.
Problem

Research questions and friction points this paper is trying to address.

Foundation Models
Explainable AI (XAI)
Image Recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interpretable AI
Image Recognition
Evaluation Methodology
🔎 Similar Papers
No similar papers found.