Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the ongoing debate regarding whether Concept Bottleneck Models (CBMs) enhance robustness by constructing a generative evaluation framework that systematically compares CBMs with standard classifiers under geometric and semantic perturbations. By disentangling robustness concepts from perturbation types and integrating randomized smoothing certification with latent space sensitivity analysis, this work proposes a controlled comparison paradigm to reconcile contradictory findings in existing literature. The results demonstrate that interpretability does not inherently confer robustness; rather, it redistributes model sensitivity, indicating that the two constitute fundamentally independent optimization objectives. Ultimately, this research clarifies the trade-off mechanisms between interpretability and robustness, providing a theoretical foundation for the design of trustworthy AI systems.
📝 Abstract
Concept Bottleneck Models (CBMs) are designed to provide interpretable intermediate representations, yet how such bottlenecks affect robustness remains unclear, with existing studies reporting mixed and sometimes contradictory findings. We argue that these discrepancies arise from conflating different robustness notions and perturbation regimes, rather than from fundamental disagreements about CBMs themselves. To disentangle these factors, we introduce a generator-based evaluation framework that enables controlled comparisons between standard classifiers and CBMs under two distinct perturbation types: continuous geometric perturbations in latent space and discrete semantic interventions in concept space. Within this framework, we evaluate robustness both empirically, via prediction and concept-level sensitivity metrics, and certifiably, using randomized smoothing in latent and concept spaces. Across experiments, we reconcile previously conflicting findings by clarifying when, and in what sense, concept bottlenecks do or do not improve robustness. By further analyzing robustness under varying task conditions, including class semantic similarity and concept vocabulary size, we show that interpretability does not inherently confer robustness. Instead, concept bottlenecks shift where and how sensitivity manifests, revealing a nuanced interpretability robustness trade off that depends critically on the perturbation regime and task structure. Together, our results show that interpretability and robustness are distinct objectives: interpretable intermediate representations do not uniformly improve robustness, but instead redistribute sensitivity across perturbation spaces and model families.
Problem

Research questions and friction points this paper is trying to address.

Concept Bottleneck Models
Robustness
Interpretability
Perturbations
Interpretability-Robustness Trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Concept Bottleneck Models
Robustness Evaluation Framework
Geometric-Semantic Perturbations
Randomized Smoothing
Interpretability-Robustness Trade-off
🔎 Similar Papers
No similar papers found.