apply concept bottlenecks

Design, build, and evaluate models whose latent representation is explicitly decomposed into human-interpretable concept variables and whose predictions are produced by first estimating those concepts (the “bottleneck”) and then mapping them to outputs; tasks include training on concept-labeled data, measuring concept prediction quality, and performing interventions or ablations on concept variables to analyze and improve model behavior and interpretability.

applyconceptbottlenecks

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Leakage and Interpretability in Concept-Based Models

Apr 18, 2025
EP
Enrico Parisini
🏛️ The Alan Turing Institute | University College London | Roche Pharmaceuticals | Kings College London

Concept bottleneck models (CBMs) suffer from pervasive information leakage—specifically, concept–task leakage (CTL), where intermediate concepts inadvertently encode task-relevant information, and inter-concept leakage (ICL), where redundant concepts exhibit spurious correlations—undermining interpretability and intervention robustness. Method: We propose the first information-theoretic, dual-metric framework to quantify CTL and ICL, formally defining and empirically validating both phenomena. We further derive theoretical links between leakage and intervention robustness, yielding actionable modeling guidelines for leakage mitigation. Contribution/Results: Our analysis reveals that CTL and ICL are prevalent across mainstream CBMs and persist irrespective of hyperparameter choices. Experiments demonstrate that our framework significantly outperforms existing methods in predicting model responses to concept interventions. It establishes a novel evaluation paradigm for trustworthy, interpretable AI and provides practical, theory-grounded optimization strategies for building robust CBMs.

Identify causes of leakage in Concept Embedding ModelsPropose guidelines to reduce leakage and ensure interpretabilityQuantify information leakage in Concept Bottleneck Models

Post-hoc Stochastic Concept Bottleneck Models

Oct 09, 2025
WJ
Wiktor Jan Hoffmann
🏛️ ETH Zurich

Pretrained concept bottleneck models (CBMs) suffer from limited modeling of inter-concept dependencies, weak causal intervention capability, and high retraining costs. To address these issues, we propose the Posterior Random Concept Bottleneck Model (PR-CBM). PR-CBM introduces a lightweight covariance prediction module—without updating the backbone network—to explicitly model the joint posterior distribution of concepts as a multivariate normal distribution, enabling efficient causal interventions. It is the first CBM variant to achieve concept-level multivariate probabilistic modeling without retraining, significantly enhancing interpretability and intervention robustness. On real-world benchmarks, PR-CBM outperforms baseline CBMs in both concept and label prediction accuracy, achieves substantially improved intervention performance, and attains markedly higher training and inference efficiency compared to fully retrained stochastic models.

Enhancing intervention performance while maintaining efficiencyImproving concept and target accuracy with minimal computationModeling concept dependencies in CBMs without full retraining

Concept Layers: Enhancing Interpretability and Intervenability via LLM Conceptualization

Feb 19, 2025
OR
Or Raphael Bidusa
🏛️ Technion - Israel Institute of Technology

Weak interpretability of large language models (LLMs) and the reliance of existing concept bottleneck models (CBMs) on manually annotated concepts and architectural modifications pose significant challenges. This paper proposes a learnable, plug-and-play Concept Layer that establishes a lightweight, differentiable mapping between internal LLM representations and an interpretable concept space—without altering the original model architecture. The layer supports both task-directed and unsupervised concept discovery and enables dynamic, inference-time intervention. Key contributions include: (i) the first zero-annotation concept set and seamless integration; (ii) ontology-guided automatic concept retrieval coupled with runtime regulation (e.g., bias mitigation); and (iii) validation of intervention efficacy via concept projection-reconstruction and interactive visualization. Experiments demonstrate that the method preserves original model performance and output consistency across multiple tasks while simultaneously enhancing interpretability, controllability, and practical utility.

Enable dynamic model interventionEnhance LLM interpretabilityIncorporate concept layers

Debugging Concept Bottleneck Models through Removal and Retraining

Sep 23, 2025
EE
Eric Enouen
🏛️ Cornell University

Concept bottleneck models (CBMs) often learn spurious concept shortcuts from biased data, leading to systematic misalignment with expert reasoning. To address this, we propose CBDebug—a novel, interpretable debugging framework that leverages concept-level expert feedback. CBDebug transforms expert judgments about undesirable concepts into sample-level auxiliary labels, enabling supervised debiasing and targeted data augmentation; it further supports model retraining after concept removal. Unlike black-box retraining approaches, CBDebug provides transparent, concept-level interventions that jointly preserve interpretability and expert alignment. Extensive experiments on multiple benchmarks exhibiting spurious correlations demonstrate that CBDebug significantly outperforms existing methods: it substantially reduces reliance on misleading concepts, improves decision consistency with human experts, and enhances model robustness.

Debugging systemic misalignment in Concept Bottleneck ModelsRemoving undesired concepts learned from biased dataRetraining models to reduce reliance on spurious correlations

Self-supervised Interpretable Concept-based Models for Text Classification

Jun 20, 2024
FD
Francesco De Santis
🏛️ Politecnico di Torino | Università della Svizzera Italiana | University of Cambridge

Weak interpretability of large language models (LLMs), unreliability of post-hoc explanation methods, and limitations of concept bottleneck models (CBMs)—including dependence on costly human annotations, restricted representational capacity, and lack of interpretability at the task level—motivate this work. We propose the self-supervised Interpretable Concept Embedding Model (ICEM), the first framework to introduce concept modeling into the textual domain. ICEM leverages the inherent generalization capability of LLMs to autonomously predict concept labels without manual annotation. It enables end-to-end interpretable prediction via concept embeddings and an interpretable decision function, supporting concept intervention, logical attribution, and controllable decoding-path steering. On text classification tasks, ICEM achieves performance comparable to fully supervised CBMs and black-box LLMs, while providing human-understandable, causally grounded explanations. Thus, ICEM unifies interpretability, interactivity, and controllability in a single architecture.

Enhancing interpretability of text analysis modelsOvercoming limitations of Concept-Bottleneck Models (CBMs)Reducing dependency on extensive concept annotations

Latest Papers

What's happening recently
View more

This work addresses the limitations of traditional Concept Bottleneck Models (CBMs), which rely on manually predefined concepts that often lack task relevance or learnability, thereby constraining model performance. To overcome this, the authors propose the Mechanistic Concept Bottleneck Model (M-CBM), which automatically extracts task-relevant, interpretable latent concepts directly from the internal representations of a black-box model using sparse autoencoders. These extracted concepts are then semantically labeled and annotated via a multimodal large language model to construct an adaptive bottleneck layer. For fair evaluation of information leakage and interpretability, the study introduces Normalized Concept Consistency (NCC), a decision-level sparsity metric. Experiments across multiple datasets demonstrate that M-CBM significantly outperforms existing CBM approaches under matched sparsity levels, achieving higher concept prediction accuracy while providing concise and human-interpretable decision rationales.

Concept Bottleneck Modelsconcept learnabilityinformation leakage

This work addresses the challenge of systematically evaluating concept bottleneck models, whose applicability and failure mechanisms remain poorly understood due to the scarcity of real-world datasets with annotated concept labels. To bridge this gap, we introduce the first controllable synthetic benchmark that leverages parametric generation techniques to precisely modulate data modality, concept selection, annotation quality, and label completeness, thereby simulating diverse real-world relationships between concepts and predictions. This benchmark enables comprehensive evaluation of various concept bottleneck models across both decision-support and fully automated tasks, effectively identifying key performance determinants and characteristic failure modes. Our framework fills a critical void in the current evaluation landscape for concept-based interpretability methods.

concept bottleneck modelsconcept labelsmodel interpretability

Existing concept bottleneck models (CBMs) struggle to verify whether their predictions rely on concepts grounded in correct visual evidence, undermining their reliability. This work proposes a fine-grained concept bottleneck model that explicitly anchors each concept to localized visual regions, enabling—for the first time—dual verification of both the existence and correctness of learned concept representations. By integrating local evidence–guided concept modeling, a verifiable architecture, and information-complete concept space learning, the proposed approach achieves prediction performance on par with standard CBMs on medical imaging benchmarks while substantially enhancing model transparency, interpretability, and the trustworthiness of concept learning.

Concept Bottleneck Modelsinterpretabilityspurious correlations

This work proposes M-CBE, a novel framework that integrates a mixture-of-experts mechanism into Concept Bottleneck Models (CBMs) to overcome the limitations of conventional approaches that rely on a single linear or Boolean predictor. By enabling multiple experts to concurrently learn diverse functional forms—such as linear models and symbolic regression—M-CBE enhances both predictive accuracy and interpretability. Moreover, it supports the automatic discovery of interpretable rules based on user-specified operator vocabularies, thereby adapting to varied interpretability requirements. Experimental results demonstrate that M-CBE effectively balances performance and explainability across multiple tasks by adjusting the number of experts and their functional forms, significantly outperforming traditional CBM methods.

accuracy-interpretability trade-offadaptabilityConcept Bottleneck Models

Hot Scholars

PB

Pietro Barbiero

SNSF Postdoctoral Fellow, IBM Research
Artificial intelligenceInterpretabilityNeurosymbolic AINeural Interpretable Reasoning
ME

Mateo Espinosa Zarlenga

University of Cambridge
Machine LearningRepresentation LearningExplainable AIInterpretability
LH

Lijie Hu

Assistant Professor, MBZUAI
Explainable AILLMDifferential Privacy
DS

Daniel Sonntag

DFKI and University of Oldenburg
Interactive Machine LearningIntelligent User InterfacesMultimodal Interaction