A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of a unified framework in existing concept-based explanation methods and the lack of rigorous theoretical guarantees for model completeness and attribution completeness. To this end, this work proposes a unified modeling framework based on a concept autoencoder. Specifically, it delineates model incompleteness through latent space analysis and reconstruction error quantification, while deriving the mathematical conditions that establish attribution completeness. For the first time, this research provides rigorous completeness bounds for concept attribution from a unified theoretical perspective. By achieving the integrated modeling of concept discovery and attribution, this work offers verifiable theoretical guarantees for interpretable artificial intelligence.
📝 Abstract
Concept-based explanations describe neural network predictions through human-understandable properties of inputs called concepts. The field encompasses approaches that differ in how they define and represent concepts and connect them to model predictions. We introduce a theoretical framework that describes these approaches in a common mathematical language and supports a shared analysis of their properties. For concept discovery, which identifies concepts automatically within a latent space of a trained model, we employ a concept autoencoder view. An encoder extracts concept representations from the model's latent space, and a decoder uses them to reconstruct the original latent representation. The autoencoder's reconstruction error measures how accurately its decoder recovers the original latent representation. We revisit model completeness: how well the concepts can reproduce the model's outputs. We show that model incompleteness of the concepts can be bounded by the autoencoder's reconstruction error. The autoencoder view also provides a common way to define individual concept attributions, which measure each concept's contribution to a prediction. We establish when these attributions sum to the model's prediction, and bound the discrepancy otherwise, thus providing attribution completeness guarantees.
Problem

Research questions and friction points this paper is trying to address.

Concept-based Explainable AI
Model Completeness
Attribution Completeness
Concept Discovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

Concept-based Explainable AI
Concept Autoencoder
Model Completeness
Attribution Completeness
Unifying Framework