Omni-Diffusion-Distill: Few-Step Distillation of Unified Multimodal Diffusion Large Language Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of iterative decoding in unified multimodal diffusion large models and the inability of existing distillation methods to jointly preserve generation and understanding capabilities. To this end, we propose the first two-stage unified distillation framework tailored for discrete token spaces. By integrating teacher trajectory replay with student self-rollback alignment, the framework simultaneously supports both modalities. Furthermore, a pairwise collision penalty is introduced to suppress text repetition, while entropy-matching guidance is designed to prevent entropy collapse during image generation. Experimental results demonstrate that our method achieves an 18.2× to 21.2× inference speedup, attaining 0.828 on GenEval and 20.0 on MM-Vet. Notably, it surpasses the performance of the teacher model under equivalent computational budgets.
📝 Abstract
Unified multimodal diffusion large language models (dLLMs) offer a single architecture for both image generation and multimodal understanding, but their iterative decoding requires tens to hundreds of forward passes. Existing few-step distillation methods largely focus on either image generation or text generation, making it unclear how to compress a fully discrete multimodal dLLM into a single efficient student while preserving both generation and understanding. We introduce Omni-Diffusion-Distill, a unified two-stage distillation framework that retains strong generation and understanding capabilities while substantially reducing the inference cost of a unified multimodal dLLM. Omni-Diffusion-Distill aligns the distillation of both generation and understanding, for both images and text, in the discrete token space. In the first stage, the student is trained to skip decoding steps by replaying cached teacher trajectories, and in the second stage the student is refined on intermediate states along its own rollouts. We further remedy two sources of degradation in unified distillation with a pairwise collision penalty that reduces repetition under parallel text decoding, and entropy-matched guidance that prevents entropy collapse caused by fitting the sharpened teacher distribution in image generation. Omni-Diffusion-Distill achieves state-of-the-art trade-offs between decoding efficiency and generation and understanding performance for multimodal dLLMs, reducing image generation from 128 to 8 decoding steps and multimodal understanding from 512 to 64, giving 18.2x and 21.2x wall-clock speedups. Under these budgets, it scores 0.828 on GenEval and 83.0 on DPG-Bench for text-to-image generation, while reaching GPT judge scores of 20.0 on MM-Vet and 57.2 on COCO captioning (twice the teacher's 28.4 at the same steps) for multimodal understanding.
Problem

Research questions and friction points this paper is trying to address.

multimodal diffusion large language models
few-step distillation
inference efficiency
unified multimodal model
discrete token space
Innovation

Methods, ideas, or system contributions that make the work stand out.

Few-Step Distillation
Unified Multimodal dLLM
Discrete Token Space
Pairwise Collision Penalty
Entropy-Matched Guidance