When Are Concept Bottleneck Model Explanations Faithful and Compact?

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of faithfulness in Concept Bottleneck Model (CBM) explanations and their inherent difficulty in simultaneously achieving compactness. We present the first formal proof that existing architectures require the inclusion of all concepts to guarantee faithfulness, thereby revealing a fundamental trade-off between these two properties. To resolve this, we propose a novel probabilistic modeling paradigm based on random variables, which integrates probabilistic graphical models with group lasso regularization to induce sparsity during training. Our approach significantly outperforms heuristic baselines in both explanation size and theoretical guarantees. By establishing rigorous conditions for CBM interpretability, this work provides a solid theoretical foundation for generating explanations that are simultaneously compact and faithful.
📝 Abstract
Concept bottleneck models (CBMs) are neural classifiers that allow to explain their decisions via high-level concepts, potentially enabling understanding, steering and debugging. However, their explanations are often derived heuristically. Building on formal explainability, we argue they should also be faithful, i.e., not misreport which concepts actually matter. We show that, for widespread CBM architectures, including recent VLM-based variants, faithful explanations must include all concepts in the bottleneck, compromising interpretability when this is large. This result applies to both heuristic and faithful-by-construction formal explanations. To encourage the existence of compact faithful explanations, we suggest i) modeling concepts probabilistically as binary or categorical random variables (rather than logits), and ii) employing per-concept training-time sparsification via group lasso (rather than regular elastic net). We also extend algorithms from formal explainability to CBMs, and show they outperform natural heuristics in terms of guarantees and explanation size. Overall, our work warns against naive interpretability claims and provides formal conditions and practical strategies for ensuring CBMs are as interpretable as advertised.
Problem

Research questions and friction points this paper is trying to address.

Concept Bottleneck Models
Faithfulness
Interpretability
Explainability
Compactness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Concept Bottleneck Models
Formal Explainability
Faithful Explanations
Group Lasso Sparsification
Probabilistic Concept Modeling
🔎 Similar Papers