🤖 AI Summary
This study addresses the challenge of safely erasing action concepts in video diffusion models by proposing MUTE, a training-free method. To our knowledge, this work presents the first systematic investigation into action erasure within video Diffusion Transformers (DiTs), establishing three essential criteria: specificity, spatial selectivity, and temporal naturalness. By conducting causal intervention analysis on attention mechanisms, MUTE extracts concept directions via token neutralization and achieves precise removal through spatial gating combined with velocity correction prior to classifier-free guidance. Experiments on Wan2.1 and CogVideoX demonstrate that MUTE significantly outperforms existing baselines, effectively suppressing target actions while preserving the natural dynamics of non-target motions.
📝 Abstract
Text-to-video (T2V) diffusion models can generate realistic depictions of actions such as kicking, stabbing, and shooting, raising safety concerns that motivate targeted concept erasure. Although concept erasure has been extensively studied for static concepts in text-to-image and T2V models, erasing motion concepts remains largely unexplored. We present a systematic study of motion concept erasure in video Diffusion Transformers (DiTs). Through causal interventions, we show that text-conditioning attention carries concept-specific motion information and supports selective intervention, whereas perturbing temporal positional encoding suppresses both target and non-target dynamics. We further find that directly adapting ESD, a representative weight-level image erasure method, to a video DiT yields modest and uneven motion suppression: reducing its erasure training loss does not by itself remove the concept signal from the difference between the conditional and unconditional predictions, which classifier-free guidance (CFG) then scales at every denoising step. From these findings, we derive three requirements for motion concept erasure: concept specificity, spatial selectivity, and temporal naturalness. Each determines one component of MUTE (Motion concept Unlearning in Text-to-video gEneration): at each denoising step, MUTE extracts a concept direction through token neutralization, derives a spatial gate from the direction's intrinsic structure, and subtracts the resulting correction from the velocity output before CFG is applied. MUTE is training-free and requires no weight modification. Experiments on 20 motion concepts show that MUTE outperforms representative prompt-level, weight-level, and inference-time baselines on Wan2.1-T2V, and the same formulation transfers to CogVideoX, supporting its applicability across distinct T2V attention architectures.