🤖 AI Summary
This study addresses the lack of effective attribution methods for analyzing generation dynamics in diffusion language models by proposing the DLIG framework. This approach extends Integrated Gradients to arbitrary network layers and denoising steps, enabling fine-grained, layer-wise and step-wise attribution analysis. By establishing a rigorous correspondence between DLIG and the axioms of Integrated Gradients, it provides a lightweight tool for verifying mechanistic hypotheses. Experiments reveal how models utilize inputs across positions, layers, and denoising steps in multi-task scenarios, quantifying their progressive commitment process during generation. This work offers a new paradigm for understanding the internal dynamics of diffusion language models.
📝 Abstract
This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-generated or fixed completion for an input prompt. We establish direct correspondences between DLIG and the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. As a lightweight complement to interventional analysis, DLIG provides an inexpensive first check of mechanistic hypotheses across the denoising trajectory. We demonstrate this on word-sense disambiguation, multi-hop graph reasoning, and sentence infilling, revealing how DLMs draw on inputs across positions, layers, and denoising steps.