🤖 AI Summary
This study addresses the vulnerability of vision foundation models to spurious correlations and the lack of interpretable, metadata-free correction methods. It proposes DiDAE, a framework integrating conditional diffusion autoencoders, sparse autoencoders, and Procrustes analysis for unsupervised dictionary learning, achieving debiasing through disentangled dictionary-based counterfactual generation. A key innovation lies in replacing gradient-based optimization with closed-form editing to enable plug-and-play deployment. Across six datasets, DiDAE outperforms state-of-the-art methods by effectively identifying and removing spurious features from classifiers, thereby significantly enhancing model robustness. Furthermore, it achieves a 2000-fold inference speedup over conventional approaches, facilitating efficient and interpretable model repair.
📝 Abstract
Foundation models remain vulnerable to spurious correlations and ``Clever Hans'' strategies. Explainable machine learning can find and remove such strategies for classifiers without metadata. For foundation models, no such option exists yet. We propose Disentangled Diffusion Autoencoders (DiDAE). DiDAE wraps a frozen foundation model in a conditional diffusion decoder. A counterfactual is one closed-form edit along a direction of a disentangled dictionary, followed by decoding. The dictionary can be supervised (Procrustes) or unsupervised (Singular Value Decomposition, Sparse Autoencoders). No gradients are needed, so DiDAE is up to 2000 times faster than the state of the art. We evaluate on six datasets, two synthetic and four real-world. In a desiderata-driven benchmark on three of them, its counterfactuals are on par with or better than the state of the art, and they repair downstream classifiers through Counterfactual Knowledge Distillation (CFKD), where they beat metadata-based correction. The same machinery can rank a pretrained dictionary against a trained classifier. It returns the few directions the classifier actually reads, each causally verified by a counterfactual that flips the decision, and repairs the classifier along those a teacher marks spurious. The workflow is plug-and-play in our open-source Peal library we publish alongside the paper. With a public dictionary and a pretrained decoder, all that remains is a cheap linear distillation of the classifier and its own fine-tuning.