Revisiting Explainable AI through Model-Independent Concept Dictionaries

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing explainable AI (XAI) methods, which often rely on model-internal abstractions or specific architectures, resulting in poor cross-model consistency and difficulty in localizing failures. To overcome these issues, this work proposes DictXAI, a novel model-agnostic concept dictionary paradigm that eliminates dependence on internal network hierarchies. By leveraging predefined dictionaries in the input domain alongside sparse coding, overcomplete dictionary learning, and attribution algorithms, DictXAI supports multimodal data—including images and waveforms—and directly maps predictions to interpretable semantic elements. Experimental results demonstrate that DictXAI effectively identifies model failures induced by data artifacts and enhances human–machine alignment in biomedical signal analysis, providing more actionable and transparent explanations than conventional approaches.
📝 Abstract
Modern applications of AI rely on increasingly complex models. Explainable AI (XAI) has emerged as a set of techniques aimed at improving model transparency. However, existing XAI methods typically assume input features to be inherently interpretable, or they rely on intermediate internal abstractions that are difficult to characterize and highly architecture-specific, hindering consistent use across models. To address these limitations, we propose DictXAI, a method that defines concepts directly in the input domain via a dictionary---a large, potentially overcomplete set of predefined elements, each carrying an interpretable meaning. Technically, DictXAI first computes a sparse code of the input and then attributes the model's prediction to the associated dictionary elements. We demonstrate the actionable nature of DictXAI explanations, showing that they can attribute AI malfunctions (e.g., Clever Hans effects) directly to identifiable artifact patterns in the data, while fostering human-AI alignment on intricate biomedical signals. We further demonstrate our method's ability to operate across a wide variety of dictionaries, including learned image bases, analytically defined waveforms for electrocardiography, and experimentally acquired dictionary elements. Overall, our results show that DictXAI provides more interpretable, actionable, and architecture-agnostic insights than classical XAI or existing concept-based approaches.
Problem

Research questions and friction points this paper is trying to address.

Explainable AI
Model-independent
Concept dictionaries
Interpretability
Architecture-agnostic
Innovation

Methods, ideas, or system contributions that make the work stand out.

Explainable AI
Concept Dictionaries
Sparse Coding
Model-Independent
Attribution