🤖 AI Summary
Large text-to-image diffusion models are often treated as black boxes, limiting artists’ understanding of and control over their generative mechanisms. This work proposes a practice-oriented interpretability approach by integrating an interactive model-bending framework into ComfyUI node-based workflows, enabling real-time intervention and visualization of specific layers in Stable Diffusion 1.5. Shifting interpretability from technical explanation toward a material perspective for creative practice, the method introduces hierarchical, feedback-driven model manipulation for the first time. Experiments demonstrate that targeted interventions on components of the diffusion pipeline consistently yield coherent visual effect families, effectively helping artists develop layered intuitions about the model’s internal mechanisms and significantly enhancing creative controllability.
📝 Abstract
Explainable AI (XAI) in creative practice can be less about technocentric explanation and more about enabling artists to inspect modify and debug models as part of making Yet largescale texttoimage diffusion systems are typically presented as opaque endtoend tools limiting this kind of material engagement We argue that even large models can function as creative materials when their internal structure is made visible and manipulable To support this we propose a handson approach to explainability centred on experimentation and intervention We instantiate this approach with a model bending and an interactive (inspection) interface integrated into ComfyUIs nodebased workflow including interactive layer selection and intervention controls Through qualitative and quantitative analysis of bending interventions in Stable Diffusion 15 we show how manipulating specific components of a diffusion pipeline produces relatively consistent families of visual effects allowing artists to build practical layerlevel intuition about how different parts of the model shape generated images