Large Language Continuous Diffusion Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of discrete diffusion models, where high-dimensional non-smooth spaces hinder inference guidance and acceleration. We propose Sigma, the first large-scale continuous language diffusion model, which constructs controllable low-dimensional ODE/SDE latent trajectories to enable parallel decoding. Methodologically, Sigma employs joint denoising and geometric learning to support embedding-space guidance and graceful degradation under few-step sampling. Training efficiency is further enhanced through block-wise likelihood optimization, autoregressive weight warm-starting, and classifier-free guidance. Empirical evaluations demonstrate that Sigma achieves performance comparable to state-of-the-art models on mathematical reasoning and code generation benchmarks, establishing a new paradigm for efficient language generation.
📝 Abstract
Despite the success of discrete diffusion language models (dLMs) for fast parallel decoding, their non-smooth, high-dimensional space hinders trajectory steering for reasoning and inference acceleration. To overcome this, we present Sigma, the first large-scale (3B/8B) continuous dLM built on steerable, low-dimensional ODE/SDE latent trajectories. Trained blockwise via likelihood optimization, Sigma jointly denoises Gaussian-corrupted token embeddings while learning an optimal embedding geometry. To accelerate training, Sigma leverages pre-trained weights from autoregressive (AR) models for warm-starting. During inference, we identify classifier-free guidance and score temperature as essential for high-fidelity reasoning and coding. Across comprehensive math reasoning and coding evaluations against state-of-the-art discrete counterparts (masked dLMs and AR baselines), Sigma achieves competitive performance with discrete models on standard benchmarks (e.g., GSM8K, Minerva, HumanEval, MBPP) after pre-training and on challenging reasoning tasks (e.g., MATH-500, AIME) after supervised fine-tuning. Beyond performance parity, we uncover key structural properties unique to continuous dLMs: (i) embedding-space steering effectively governs the quality-diversity trade-off, yielding strong pass@k performance and (ii) continuous trajectories enable graceful degradation for low NFEs and efficient distillation. These establish continuous dLMs as a promising paradigm for efficient language generation.
Problem

Research questions and friction points this paper is trying to address.

discrete diffusion language models
continuous diffusion
trajectory steering
reasoning
inference acceleration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continuous Diffusion Language Models
Latent Trajectories
Classifier-Free Guidance
Embedding-Space Steering
Efficient Distillation