🤖 AI Summary
This study addresses the disconnect between sparse feature interpretation and generative control in diffusion models, along with the absence of annotation-free retrieval and intervention verification mechanisms. To this end, we propose D-Scope, a framework that leverages sparse autoencoders for feature extraction and achieves annotation-free contrastive retrieval by matching SigLIP visual centroids with textual descriptions. Furthermore, it employs spatial mask interventions to verify how decoder directions influence generation, thereby establishing a shared visual evidence chain linking feature interpretation and generative control. Our experiments reveal the coexistence of high reconstruction fidelity and low dictionary utilization, and demonstrate that contrastive retrieval significantly outperforms direct retrieval in terms of regional gain. Ultimately, this work provides an evaluable, comprehensive paradigm for the interpretability of diffusion models.
📝 Abstract
Sparse autoencoders (SAEs) reveal visual structure in diffusion transformers (DiTs), but interpreting a feature does not establish whether it can be used to control generation. We introduce D-Scope (Diffusion Scope), a framework that connects feature interpretation to generation control through shared visual evidence. D-Scope aggregates SigLIP~2 embeddings of highly activating image patches into visual centroids. Matching target text descriptions against these visual centroids in the shared image-text embedding space then enables retrieval of individual features without per-feature text annotations. The underlying patches provide evidence for inspecting each selection, while spatially masked interventions test the corresponding decoder direction at varying strengths under fixed generation conditions. We characterize 150 SAEs across two model families and five layers, and introduce a benchmark of 100 target concepts with ten contexts each spanning under-specified and explicit-conflict conditions. Our empirical results show that high reconstruction fidelity can coexist with low dictionary utilization and limited visual-evidence coverage. Under per-case best-of-sweep strength selection, contrastive retrieval yields larger mean regional SigLIP~2 gains than direct retrieval across the tested steering configurations, without consistently improving outside-region preservation. D-Scope provides an inspectable framework for evaluating sparse DiT features through their visual evidence and the effects of their decoder directions on generation. The demo is available at https://jiahaozhang-public.github.io/d-scope/.