E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of masked diffusion models, where the positional factorization assumption in the reverse process degrades few-step generation quality and continuous latent variables are prone to posterior collapse. To overcome these issues, this work proposes a discrete shared latent variable approach based on a Mixture-of-Experts (MoE) architecture. By introducing cross-positional dependencies through expert routing decisions, the method achieves non-factorized modeling without increasing the number of active parameters. Furthermore, it reformulates the reverse process as a mixture distribution to effectively circumvent posterior collapse. Experimental results demonstrate that the proposed approach significantly improves few-step generation performance across synthetic data, MNIST, and the LM1B benchmark.
📝 Abstract
Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of work introduces a continuous Gaussian latent, trained as a variational autoencoder, to capture correlations across positions, but such approaches are prone to posterior collapse, where the latent is silently ignored. We propose Enhanced Mixture-of-Experts (E-MoE), which builds the reverse process as a mixture of factorized distributions over a discrete shared latent given by the expert-routing decisions of a Mixture-of-Experts (MoE) backbone, without increasing active parameters over the factorized baseline. Across synthetic multi-modal benchmarks, binarized MNIST, and LM1B, E-MoE improves few-step generation over factorized baselines.
Problem

Research questions and friction points this paper is trying to address.

Masked diffusion models
Few-step generation
Factorized reverse process
Posterior collapse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Masked Diffusion Models
Non-Factorized Reverse Process
Discrete Shared Latent
Few-step Generation
🔎 Similar Papers
No similar papers found.