Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the opacity of semantic consistency mechanisms in text-to-image diffusion Transformers (DiTs) during denoising by introducing a causal interpretability framework that integrates attention decomposition, directed interventions across tokens, heads, and layers, and token-level ablation. The study reveals— for the first time—that object identity is preserved not by input semantic tokens, but by structural template tokens acting as implicit semantic registers. It further elucidates the staged computational organization underlying semantic routing and visual synthesis in DiTs and proposes a training-agnostic attention head pruning strategy. This approach reduces attention FLOPs by 20% with only a 1.4-point drop in GenEval performance, offering a systematic dissection of the generative pipeline from identity formation and propagation to refinement.
📝 Abstract
Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that combines attention decomposition with targeted interventions across token spans, heads, and layers. Using it to separate prompt-content tokens from structural template tokens, we find that the structural tokens carry little prompt-specific information at the encoder output. Yet surprisingly, they emerge as dominant image-to-text attention sinks and causally maintain object identity inside the DiT, acting as implicit semantic registers. We show that they acquire this identity indirectly, with prompt semantics first injected into the image latents and then read back into the template tokens rather than transferred directly from the prompt tokens. Inspired by the above findings, we design a training-free pruning rule for DiTs. Heads that attend most strongly to prompt tokens are dispensable, and pruning them removes $20\%$ of attention FLOPs with only a $1.4$-point drop on GenEval. We further reveal how generative computation in DiTs is organized across heads and depth, separating semantic routing from visual synthesis and progressing from identity formation to propagation and refinement. Our work not only reveals that the tokens encoding semantics at input need not be those that maintain it during generation, but also provides a causal view of internal mechanisms in DiTs.
Problem

Research questions and friction points this paper is trying to address.

diffusion transformers
text-to-image generation
semantic representation
token dynamics
interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

diffusion transformers
causal interpretability
semantic registers
attention pruning
text-to-image generation
Maohua Li
Maohua Li
Hohai University
Spiking Neural Networks
Q
Qirui Li
Alibaba Group, Zhejiang University
Y
Yanke Zhou
Nanjing University
Y
Yiduo Li
Alibaba Group
Z
Zhaosheng Chi
Alibaba Group
C
Chao Xu
Alibaba Group
C
Cuifeng Shen
Alibaba Group
Y
Yixuan Xu
Alibaba Group
H
Hanlin Tang
Alibaba Group
K
Kan Liu
Alibaba Group
T
Tao Lan
Alibaba Group
L
Lin Qu
Alibaba Group
S
Shao-Qun Zhang
Nanjing University