🤖 AI Summary
This study addresses the limitation of existing CLIP explanation methods in revealing how input modifications can achieve target predictions. To this end, we propose MACE, a framework that leverages adaptive masking to guide latent diffusion models in editing image regions, thereby generating targeted counterfactual explanations for CLIP classification. The core innovation lies in introducing the first adaptive mask expansion mechanism, which effectively balances the magnitude of pixel alterations against target achievement rates and elucidates the trade-off inherent in preserving validity. Experimental results demonstrate that the proposed method achieves state-of-the-art target success rates and visual realism across four datasets, significantly outperforming Stable Diffusion baselines.
📝 Abstract
Current explanation methods for contrastive vision--language models such as CLIP mainly identify important regions without showing how to change the input in order to get a target prediction. We introduce \textbf{M}ask-guided \textbf{A}daptive \textbf{C}ounterfactual \textbf{E}xplanations (\mace), a targeted visual counterfactual method designed specifically for CLIP zero-shot classification. \mace constructs an editable region from either source attribution or source--target attribution differences and expands the mask only when needed to reach a specified target class. A latent diffusion inpainting model then modifies the selected region, while a frozen CLIP model provides modification guidance and anchors the remaining image content to the original input. We evaluate \mace on ImageNet, Food-101, Oxford Pets, and CUB-200. The source-mask variant achieves the highest target top-1 success rate across all four datasets, while the difference-mask variant produces the smallest pixel-level and perceptual changes and the best realism scores. Both variants improve proximity and realism over a Stable Diffusion-only baseline using the same generative backbone. These results show that adaptive mask-guided editing produces effective CLIP counterfactuals. They further reveal a tradeoff between counterfactual validity and source-image preservation.