🤖 AI Summary
This study addresses the inefficiency and knowledge forgetting challenges encountered when converting autoregressive models to diffusion models under low computational budgets. Specifically targeting Mixture-of-Experts (MoE) large language models, this work proposes a frozen context tower mechanism that integrates cross-attention, denoising loss, and representation alignment techniques. This approach enables efficient architectural conversion with limited data, significantly outperforming conventional in-place fine-tuning strategies. The proposed method effectively preserves the generative capabilities and prior knowledge of the parent model, achieving an 11.6-fold improvement on HumanEval while retaining 95% performance on GSM8K and 99% on MMLU-Pro. Consequently, this research establishes a novel paradigm for low-cost model conversion.
📝 Abstract
Converting a pretrained autoregressive (AR) model to a diffusion language model (dLLM) enables parallel generation without pretraining a new model. Published conversion methods differ by roughly three orders of magnitude in training data and have not been compared under a common protocol. We compare two conversions of the same 30B Mixture-of-Experts (MoE) parent, holding the corpus, supervised-token budget, trainable parameter set and evaluation harness fixed, each under its own training recipe. The in-place model updates a subset of the parent's weights using denoising and representation-alignment losses; the frozen-tower model instead conditions through cross-attention on a frozen causal copy of the parent. With 1B training tokens, the frozen-tower model scores 71.60 on HumanEval pass@10 against 6.19 for the in-place model, an 11.6x improvement. At the same budget it also keeps 95% of the parent's GSM8K score and 99% of its MMLU-Pro score. A dense-parent experiment reproduces the HumanEval separation. Within the two-tower design at about 500M tokens, freezing the context tower retains substantially more MMLU-Pro performance than training it, while both give similar observed HumanEval scores. Our theoretical analysis establishes that both conversion classes contain an exact sampler for the AR parent under a hard attention mask and left-to-right commitment of one position per round. Under a shared loss, freezing removes the gradient contribution through the context states. Furthermore, evaluation protocol substantially affects a published 500B-token conversion's scores in both directions across tasks, while its AR parent's scores vary by less than three points, so comparing dLLMs needs a common protocol. These results show that, in the tested low-budget regime, the frozen-tower configuration retains substantially more of the parent's generation performance than in-place conversion.