🤖 AI Summary
This study addresses the limited generalizability of existing backdoor defenses for text-to-image models, which struggle against diverse attack mechanisms. To this end, this work proposes the NDDL framework, which leverages the transition dynamics of the diffusion process to detect and localize backdoor triggers by learning multi-space trajectory evolution patterns exclusively from benign samples. Specifically, the framework integrates timestep-conditioned dynamics modeling, multi-space compact representation, and low-semantic word substitution localization. Without requiring any prior knowledge of the attack, NDDL achieves efficient, generalized defense against various unknown backdoor attacks while enabling precise trigger localization.
📝 Abstract
Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack mechanisms. In this paper, we study backdoor defense of T2I diffusion models from a transition-dynamics perspective. We observe that benign diffusion trajectories exhibit structured and timestep-dependent transition patterns from cross-attention, latent and noise spaces, whereas backdoor attacks tend to induce deviations from such normal evolution. Motivated by these observations, we propose Normal Diffusion Dynamics Learning (NDDL), a novel backdoor defense framework that learns the normal transition dynamics of diffusion trajectories utilizing only benign samples. NDDL constructs compact multi-space trajectory representations and trains a timestep-conditioned dynamics model to predict the diffusion evolution. In the inference phase, deviations between the observed and predicted transitions are exploited to quantify dynamics inconsistency for backdoor detection. NDDL further enables trigger localization without any prior knowledge of the embedded backdoor by performing substitution with low-semantic words. Extensive experiments for diverse backdoor attacks demonstrate the effectiveness and generalizability of our proposed NDDL.