ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance gap between continuous diffusion language models and autoregressive models in mathematical reasoning and code generation. We propose REPA+REG, a method grounded in embedded language flows and continuous diffusion modeling. By introducing representation alignment and entanglement mechanisms, it leverages a frozen autoregressive teacher model to supervise intermediate feature learning, enabling joint denoising of global representations and supporting early-stopping decoding strategies without few-step training. Experimental results demonstrate that the proposed approach achieves 55.96% accuracy on GSM8K, significantly outperforming baselines of comparable scale. Furthermore, it attains a pass@10 of 41.21% on HumanEval with only 16 neural function evaluations (NFEs), effectively enhancing the complex reasoning capabilities of diffusion language models.
📝 Abstract
Fully continuous diffusion language models (dLMs) denoise continuous representations without intermediate discretization, then decode all response tokens in parallel at the final step. Their performance on challenging reasoning tasks remains less established than that of autoregressive (AR) LLMs and masked dLMs. We scale Embedded Language Flows (ELF) to mathematical reasoning and code generation on GSM8K, MATH-500, HumanEval, and MBPP. We introduce ELF-REG, which improves learning with representation alignment and entanglement (REPA+REG), where a frozen AR teacher supervises intermediate denoiser features and supplies a global representation that is jointly denoised with the response. ELF-REG-L achieves 55.96% pass@1 on GSM8K at 64 network function evaluations (NFE), and 13.39% on MATH-500 and 22.56% on HumanEval at 128 NFE. It outperforms the evaluated comparable-scale dLMs in pass@1 on GSM8K and code, and improves MATH-500 pass@1 from 10.55% for the ELF-L baseline to 13.39% with ELF-REG-L. Without few-step training, the same task-specific checkpoints support strong low-NFE performance through early-stop, which decodes an intermediate clean prediction without completing the denoising trajectory. At 16 NFE, ELF-REG-L reaches 41.21% HumanEval pass@10, outperforming recent continuous dLMs of comparable scale.
Problem

Research questions and friction points this paper is trying to address.

continuous diffusion language models
reasoning tasks
mathematical reasoning
code generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continuous Diffusion Language Models
Representation Alignment and Entanglement
Reasoning Tasks
Early-stop Decoding
Teacher Supervision
Z
Zeyu Michael Li
Duke University
W
William Xingxu Chen
Duke University
B
Bingshuo Qian
Duke University
J
Jiayin Liu
Tsinghua University
Xiang Cheng
Xiang Cheng
Duke University
stochastic processesmachine learningtheory of deep learning