Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations
This study addresses the ongoing debate regarding whether Concept Bottleneck Models (CBMs) enhance robustness by constructing a generative evaluation framework that systematically compares CBMs with standard classifiers under geometric and semantic perturbations. By disentangling robustness concepts from perturbation types and integrating randomized smoothing certification with latent space sensitivity analysis, this work proposes a controlled comparison paradigm to reconcile contradictory findings in existing literature. The results demonstrate that interpretability does not inherently confer robustness; rather, it redistributes model sensitivity, indicating that the two constitute fundamentally independent optimization objectives. Ultimately, this research clarifies the trade-off mechanisms between interpretability and robustness, providing a theoretical foundation for the design of trustworthy AI systems.