๐ค AI Summary
The mechanisms underlying the emergence of semantic structure during diffusion model generation remain poorly understood, and existing approaches struggle to simultaneously capture the dynamic evolution of attention across both spatial and temporal dimensions. This work proposes a novel visual analytics framework that, for the first time, integrates timestep-indexed token-level cross-attention maps with data-driven phase identification. By combining time-series clustering, quantitative attention metrics, and interactive visualization, the framework enables structured analysis of attention dynamics in Stable Diffusionโlike models. Evaluated on a benchmark of 60 structured prompts, it reveals interpretable patterns of attention evolution, effectively supporting human-in-the-loop understanding and control of the generative process.
๐ Abstract
Diffusion-based text-to-image models can synthesize complex and highly structured visual content, yet the emergence and evolution of semantic structure remain difficult to interpret. Many existing workflows rely on aggregated attention or scalar summaries that separate temporal change from image-space evidence. To address this gap, we present a visual analytics framework for exploring attention dynamics in diffusion models: the step-indexed evolution of token-level cross-attention maps, their temporal concentration, and their spatial relationships. Our approach enables structured analysis of attention behavior across generation steps by integrating quantitative measures with data-driven stage identification in an interactive workflow. Case studies on a structured 60-prompt Stable-Diffusion-class benchmark illustrate recurring, interpretable patterns within this setting and show how linked temporal and spatial views facilitate the observation and discussion of generative processes, supporting more effective human-AI collaboration.