DIET: Deletion-response Expert Trimming for Video Diffusion Transformers

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high storage costs of Mixture-of-Experts (MoE) architectures in video diffusion Transformers and the limitation of existing pruning methods that overlook inter-layer rerouting effects. We propose a training-free expert pruning framework that introduces a novel deletion response signature mechanism via cached tensor replay, enabling precise quantification of expert removal impacts on conditional and unconditional tokens without additional forward passes. Furthermore, retained budgets are allocated through an optimized diversity loss combined with a hierarchical search strategy integrating intra-layer local search and inter-layer regression guidance. Evaluated on LingBot-Video, our method achieves 50% expert pruning, compressing the model from 57GB to 30GB to enable single-GPU deployment while improving the overall VBench score to 0.8115.
πŸ“ Abstract
Video diffusion transformers (DiTs) increasingly adopt mixture-of-experts (MoE) architectures to reduce active computation, but their full expert storage remains costly. Existing one-shot pruning criteria mainly rely on static activation or routing statistics and cannot capture layer-level re-routing after expert deletion. We introduce DIET, a training-free expert pruning framework based on deletion responses. A single all-expert calibration pass records expert outputs and router states for matched conditional and unconditional tokens. Candidate deletions are then replayed from cached tensors, requiring no additional model forward passes. The resulting deletion-response signatures characterize each expert by the changes induced when it is removed. DIET selects retained experts by minimizing Overall Diversity Loss (ODL), which preserves directional coverage in signature space, and combines intra-layer local search with an inter-layer regression-guided budget search to allocate experts across layers. On LingBot-Video 30B-A3B, pruning 50% of experts (6,144 to 3,072) reduces the checkpoint from 57 GB to 30 GB and enables single-card deployment on a 48 GB GPU without fine-tuning. Under a fixed 284-case VBench protocol, the VBench Total increases from 0.7941 to 0.8115. Across tested retention budgets, DIET consistently outperforms competitive pruning baselines adapted from large language models.
Problem

Research questions and friction points this paper is trying to address.

Video Diffusion Transformers
Mixture-of-Experts
Expert Pruning
Model Compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Diffusion Transformers
Mixture-of-Experts Pruning
Training-free Expert Trimming
Deletion-response Signatures
Overall Diversity Loss
πŸ”Ž Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30