ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational and storage bottlenecks arising from the massive expert parameters in Mixture-of-Experts (MoE) diffusion language models by proposing an importance-guided, token-aware compression framework. The proposed method integrates adaptive Tucker decomposition with token-aware compensatory routing, dynamically allocating low-rank dimensions through activation and gradient importance estimation while differentially processing hot and cold tokens to eliminate cross-modal redundancy. Experimental results demonstrate that the framework retains 96.33% accuracy at a 30% compression ratio and achieves up to a 7.22× end-to-end inference speedup, striking an excellent balance between model efficiency and generative performance.
📝 Abstract
Mixture-of-Experts (MoE) Diffusion Language Models (DLMs) offer flexible parallel decoding and increased model capacity, but their large number of expert parameters incurs substantial computation and storage costs. Existing low-rank MoE compression methods largely rely on static factorization and fixed rank allocation, which overlook the distinctive properties of MoE DLMs. Specifically, we identify two properties: cross-mode non-uniform redundancy, where parameter redundancy and sensitivity to rank truncation vary across the input, output, and expert modes, and token-wise utilization variation, where hot and cold tokens exhibit distinct spectral characteristics and expert activation patterns. To address these challenges, we propose ITC-MoE, an Importance-guided Token-aware Compression framework for MoE DLMs. ITC-MoE consists of two complementary components. First, Importance-guided Adaptive Tucker Compression (IATC) incorporates activation and gradient importance into expert weight transformation, jointly factorizes expert weights across multiple modes, and adaptively allocates ranks under a fixed parameter budget. Second, Token-aware Compensation and Routing (TCR) applies lightweight low-rank compensation to compression-sensitive hot tokens and restricts the candidate expert set for cold tokens with concentrated routing patterns. By jointly adapting compression capacity and inference execution to both parameter redundancy and token-wise variation, ITC-MoE substantially reduces the computation and storage costs of MoE DLMs while preserving their generation quality. For example, on SDAR-30B-A3B-Chat-b32, ITC-MoE maintains an accuracy of 96.33% on MultiArith under a 30% compression budget, while achieving up to a 7.22x end-to-end speedup. The code is publicly available at https://github.com/lianjunl13-sudo/ITC-MoE.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Diffusion Language Models
Model Compression
Parameter Redundancy
Token-aware
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Diffusion Language Models
Adaptive Tucker Compression
Token-aware Routing
Model Compression
💼 Related Jobs
No related jobs found.
L
Lianjun Liu
Hainan University, Haikou, China
S
Shipeng Li
Hainan University, Haikou, China
You Huang
You Huang
Xiamen University
segmentationinteractive segmentationtransformer
W
Weiqi Yan
Xiamen University, Xiamen, China
M
Mingte Qiu
Hainan University, Haikou, China
Huazhong Liu
Huazhong Liu
Huazhong University of Science and Technology
computer sciencebig datahigh performance computing
X
Xiaofeng Zhu
Hainan University, Haikou, China
Yunshan Zhong
Yunshan Zhong
Hainan university