Score
Designs and builds concept bottleneck model architectures that factor concept representations by object parts, including part-factorized CBMs and part-based bottlenecks that route attribute information to part-specific tokens and spatially ground concept predictions to localized parts. Develops training and supervision methods to produce concept predictions before class decisions for auditability and evaluates approaches that preserve overall classification accuracy with minimal part-level supervision.
Concept Bottleneck Models (CBMs) predict predefined concepts but fail to satisfy the information bottleneck principle—concept prediction capability does not imply concept-exclusive encoding, undermining interpretability and intervention reliability. Method: We identify this fundamental limitation and propose the Minimum Concept Bottleneck Model (MCBM), which enforces each latent variable to retain only the minimal sufficient information for its associated concept via variational information bottleneck regularization. Contribution/Results: MCBM is the first CBM framework to provide theoretical guarantees for concept interventions, ensuring both Bayesian consistency and architectural flexibility. Empirical evaluation across multiple benchmarks demonstrates significant improvements in concept specificity, intervention robustness, and model interpretability. These results validate the critical role of strict information constraints in building trustworthy, concept-based models.
This work addresses the unreliability of concept–prediction associations in Concept Bottleneck Models (CBMs), often caused by concept leakage and accompanied by degraded accuracy. The authors propose a theoretically grounded, architecture-agnostic information bottleneck regularization method that learns minimal sufficient concept representations by minimizing the mutual information \(I(X;C)\) between inputs and concepts while preserving the mutual information \(I(C;Y)\) between concepts and labels—all without modifying model architecture or requiring additional supervision. By integrating a variational objective with entropy-based proxy constraints, the approach seamlessly fits into standard CBM training pipelines. Information plane analysis confirms its mechanistic efficacy. Evaluated across six CBM variants and three benchmark datasets, the method consistently improves prediction accuracy, mitigates concept leakage, and enhances the stability of concept interventions.
This work investigates whether Concept Bottleneck Models (CBMs) satisfy the locality assumption—that concept predictions depend solely on features genuinely relevant to each concept, rather than spurious, statistically confounded features. Method: We conduct a systematic analysis via input perturbation, causally inspired concept attribution diagnostics, theoretical modeling, and empirical evaluation across multiple benchmarks. Contribution/Results: We are the first to demonstrate that CBMs fundamentally violate locality—even under ideal conditions of concept independence and non-overlapping features—due to implicit inter-concept correlations that induce interpretability fragility. Empirically, CBMs frequently exploit spurious features to achieve high accuracy, resulting in hollow concept explanations. Locality violation is thus identified as the root cause of their compromised interpretability and robustness. Our findings expose a critical flaw in the foundational assumptions underlying concept-based interpretability and provide both theoretical warnings and practical diagnostic tools for building trustworthy concept models.
Existing concept bottleneck models (CBMs) struggle to capture high-order interactions among concepts and cannot quantify the conditional dependence probabilities between concepts and predictions, limiting their interpretability and causal intervention capability. To address these limitations, we propose the Energy-driven Concept Bottleneck Model (ECBM), the first CBM that formulates the concept bottleneck as a differentiable joint energy function over inputs, concepts, and labels. This unified representation explicitly encodes nonlinear concept interactions and cross-level conditional dependencies, thereby overcoming the restrictive independence assumptions and intervention failures inherent in conventional CBMs. Our method integrates neural energy function design, energy decomposition–based conditional probability derivation, contrastive learning, and MCMC-based approximate inference. Evaluated on multiple real-world datasets, ECBM achieves significant improvements over state-of-the-art methods in classification accuracy, while enabling computable concept correction paths and fine-grained probabilistic explanations.
Existing concept bottleneck models (CBMs) heavily rely on high-quality, labor-intensive human concept annotations and frequently suffer from misalignment between concept saliency and input saliency. To address the challenge of scarce labeled data, this paper proposes a semi-supervised concept bottleneck model (SSCBM)—the first to integrate semi-supervised learning into the CBM framework. SSCBM introduces a concept-level pseudo-labeling strategy and a concept-space alignment loss to enforce consistency constraints at the concept level over unlabeled samples. By jointly optimizing on both labeled and unlabeled data, SSCBM achieves 93.19% concept accuracy and 75.51% prediction accuracy using only 20% of the labeled data, approaching fully supervised performance (96.39% / 79.82%). This significantly reduces expert annotation effort while mitigating concept–input misalignment.
This work addresses the high intervention cost and reliance on extensive expert annotations inherent in traditional concept bottleneck models, which also require maintaining multiple separate models. The authors propose a unified architecture that, for the first time, integrates Matryoshka representation learning into concept bottleneck models. By constructing a nested, hierarchical concept structure guided by the principle of maximum relevance and minimum redundancy, the approach enables multi-granularity adaptive inference within a single model. Crucially, it allows dynamic adjustment of concept usage during inference without retraining, reducing intervention complexity from linear to logarithmic while ensuring monotonic performance improvement. The method achieves accuracy comparable to that of independent models yet substantially lowers expert annotation overhead, thereby facilitating efficient and flexible human-AI collaborative reasoning.
This work addresses the issue of information leakage in existing Concept Bottleneck Models (CBMs), which occurs when the number of concepts approaches the embedding dimension, leading models to rely on spurious correlations and compromising interpretability. To mitigate this, the authors propose the Concept Flow Model (CFM), a hierarchical, concept-driven differentiable probabilistic decision tree that focuses on locally discriminative concepts at each internal node to progressively narrow down predictions. CFM leverages vision-language models to generate concept embeddings and constructs a decision hierarchy grounded in visual embeddings, optimizing hierarchical concept weights in an end-to-end manner. Experiments demonstrate that CFM achieves prediction performance comparable to flat CBMs while significantly reducing the number of effective concepts used, thereby alleviating information leakage and enabling transparent, auditable reasoning aligned with hierarchical class structures.
This work addresses the challenge that Concept Bottleneck Models (CBMs) rely on high-quality concept annotations, which are scarce in practice, while concepts generated solely by Vision-Language Models (VLMs) often lack sufficient fidelity, undermining model interpretability. To overcome this limitation, the authors propose VH-CBM, a novel approach that effectively integrates VLM-derived priors with minimal human-annotated dense labels. By leveraging Gaussian processes to propagate expert supervision within the VLM embedding space and incorporating active learning to optimize annotation efficiency, VH-CBM achieves substantial improvements in both concept prediction accuracy and calibration. Experimental results demonstrate that with only 1% of the annotated data, VH-CBM significantly outperforms purely VLM-guided CBM variants.
研究通过大规模用户实验评估了概念瓶颈模型(CBMs)作为决策支持系统的有效性,发现其在特定条件下能提高人机团队的准确性。
Existing concept bottleneck models (CBMs) struggle to verify whether their predictions rely on concepts grounded in correct visual evidence, undermining their reliability. This work proposes a fine-grained concept bottleneck model that explicitly anchors each concept to localized visual regions, enabling—for the first time—dual verification of both the existence and correctness of learned concept representations. By integrating local evidence–guided concept modeling, a verifiable architecture, and information-complete concept space learning, the proposed approach achieves prediction performance on par with standard CBMs on medical imaging benchmarks while substantially enhancing model transparency, interpretability, and the trustworthiness of concept learning.