LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study systematically investigates the scaling laws of Mixture-of-Experts (MoE) diffusion language models, addressing critical theoretical gaps in optimization, data-to-model allocation, and architectural design. It reveals, for the first time, quantitative differences in scaling behavior between MoE diffusion models and autoregressive counterparts, and introduces IsoFLOP—a computation allocation strategy tailored for diffusion language models—alongside principles for stable expert architecture design. Leveraging these insights, the authors train LLaDA MoE v2 from scratch, featuring 30 billion total parameters with only 3 billion activated per token, using merely 65% of Qwen3’s pretraining data (23.5T tokens). The model matches Qwen3’s performance across diverse knowledge, reasoning, and code tasks; after supervised fine-tuning, it surpasses SDAR Chat on seven out of eight benchmarks.
📝 Abstract
Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood. We systematically characterize how optimization hyperparameters, compute allocation, and architecture scale for MoE dLLMs, identifying quantitative differences from scaling trends previously reported for AR models. Specifically, for optimization, the optimal nominal batch size grows faster, while the optimal learning rate decays more rapidly with compute. For model--data allocation, IsoFLOP analysis reveals a slight data-side tilt: the optimal token budget grows faster than activated model-side computation. For MoE architecture, larger scales increasingly favor larger expert pools at fixed activated capacity, while moderate expert granularity remains consistently effective and the preferred fraction of activated capacity assigned to shared experts remains stable across scales. Guided by these findings, we train LLaDA MoE v2, a 30B-A3B dLLM, from scratch on 23.5T tokens. With approximately 65\% as many pretraining tokens as Qwen3, LLaDA MoE v2 approaches Qwen3 on several knowledge, reasoning, and coding benchmarks. After supervised fine-tuning alone, it outperforms SDAR Chat on seven of eight reasoning and coding benchmarks and remains close to Qwen3 on several tasks. These results establish practical scaling laws and design principles for MoE dLLMs.
Problem

Research questions and friction points this paper is trying to address.

diffusion language models
Mixture-of-Experts
scaling laws
compute allocation
model architecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
diffusion language models
scaling laws
IsoFLOP analysis
expert architecture
🔎 Similar Papers
F
Fengqi Zhu
Gaoling School of AI, Renmin University of China; Beijing Key Laboratory of Research on Large Models and Intelligent Governance; Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE
S
Shaoxuan Xu
Gaoling School of AI, Renmin University of China; Beijing Key Laboratory of Research on Large Models and Intelligent Governance; Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE
J
Jingyang Ou
Gaoling School of AI, Renmin University of China; Beijing Key Laboratory of Research on Large Models and Intelligent Governance; Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE
Zebin You
Zebin You
renmin university of china
generative modeldiffusion modelsemi-supervised learningself-supervised learning
Y
Yipeng Xing
Ant Group
Huabin Liu
Huabin Liu
Shanghai Jiao Tong University
Computer VisionVideo UnderstandingVideo ReasoningFew-shot Learning
Xiaolu Zhang
Xiaolu Zhang
Staff Algorithm Engineer, Ant Group
Deep learningComputer VisionRecommendation SystemMedical Imaging
Jun Zhou
Jun Zhou
Ant Group, Alibaba Group, Zhejiang University
Distributed Machine LearningPrivacy Preserving Machine LearningGraph Neural NetworksAutoML
Zhenzhong Lan
Zhenzhong Lan
School of Engineering, Westlake University
NLPComputer VisionMultimedia
Yankai Lin
Yankai Lin
Associate Professor (Tenure Track), Gaoling School of AI, Renmin University of China
Natural Language ProcessingLarge Language Models
Wayne Xin Zhao
Wayne Xin Zhao
Professor, Renmin University of China
Recommender SystemNatural Language ProcessingLarge Language Model
Jianguo Li
Jianguo Li
Director, Ant Group
deep learningcomputer visionmachine learningsystem
Chongxuan Li
Chongxuan Li
Associate Professor, Renmin University of China
Machine LearningGenerative ModelsDeep Learning
Ji-Rong Wen
Ji-Rong Wen
Gaoling School of Artificial Intelligence, Renmin University of China
Large Language ModelWeb SearchInformation RetrievalMachine Learning