Hierarchical Continuous Diffusion Language Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing diffusion language models, where discrete variants lose token dependencies during parallel decoding and continuous counterparts lack effective token-level constraints. To overcome these issues, this work proposes a hierarchical continuous diffusion language model that couples discrete generation with continuous latent trajectories within a unified denoising process. By establishing the latent variable as the core state, tokens are decoded at each step and fed back as scaffolding for subsequent updates, thereby circumventing the limitations of independent sampling. The training objective is derived from a variational bound on token likelihood, enabling bidirectional inference. Empirical evaluations demonstrate that the proposed approach significantly outperforms existing discrete and continuous baselines in accuracy on Sudoku and Countdown puzzles, as well as in perplexity on the LM1B benchmark.
📝 Abstract
Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update. On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B. Project page: https://hc-dlm.github.io/.
Problem

Research questions and friction points this paper is trying to address.

diffusion language models
discrete diffusion
continuous diffusion
parallel decoding
statistical dependencies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continuous Diffusion
Hierarchical Latent
Variational Bound
Language Models
Non-autoregressive Generation
🔎 Similar Papers
2022-09-02ACM Computing SurveysCitations: 1628