🤖 AI Summary
This study addresses the issue of unfaithful representations caused by correlated concept entanglement in Concept Bottleneck Models by proposing an interpretable layer based on the 2-additive Choquet integral. By integrating the CLIP vision-language model with closed-form gradient derivation, this method drives concept organization without explicit supervision, merging correlated concepts into compact nodes whose weights directly map to Shapley values to support test-time interventions. Experiments across four datasets demonstrate that the proposed approach generates sparse and semantically coherent nodes, effectively optimizing the accuracy-interpretability trade-off while achieving debiasing performance comparable to methods requiring retraining.
📝 Abstract
Concept Bottleneck Models (CBMs) built on vision-language models such as CLIP represent a latent space as human-understandable concepts. These representations are unfaithful: related concepts are entangled, so individual scores do not reflect their intended meaning. We propose CHOQOLATE, an interpretable-by-design layer based on 2-additive Choquet integrals, which merges correlated concepts into compact nodes. Across four datasets, CHOQOLATE achieves a favorable accuracy-interpretability trade-off, with weight-sparse and semantically coherent nodes. A closed-form gradient derivation, backed by experiments, explains why Choquet layers drive this organization without explicit supervision. Choquet weights also map directly to Shapley values, which enables test-time intervention. On standard bias-mitigation benchmarks, suppressing spurious concepts after training performs on par with methods that require group annotations or retraining, while needing neither.