🤖 AI Summary
This study addresses the high generation costs of continuous diffusion language models and the challenges of modeling and decoding compressed embeddings in two-stage approaches. To overcome these limitations, this work proposes JPEG-DLM, which adopts a flow matching framework and employs a joint embedding prediction mechanism to enable end-to-end training of the compressor, flow matching model, and decoder. This design learns structured and easily decodable compressed embeddings, effectively circumventing fixed-space constraints. Experimental results demonstrate that JPEG-DLM achieves the lowest generation perplexity and highest throughput on the LM1B and OWT datasets. Notably, at a 0.5 compression ratio, it attains a generation perplexity (Gen-PPL) of 34.52 with a throughput 2.3 times that of ELF.
📝 Abstract
Most continuous diffusion language models process one latent position per token at each sampling step, making generation expensive. Two-stage methods lower the cost by reducing the latent length, but they fix the compressed embedding space before training the diffusion model. Embeddings from the fixed space can be difficult to model with diffusion and decode reliably into tokens, which limits generation quality after compression. To address this problem, we introduce JPEG-DLM (Joint-embedding Prediction for Efficient Generation with Diffusion Language Model), which jointly trains a compressor, a flow matching model and a decoding module. With joint-embedding prediction, JPEG-DLM learns compressed embeddings that are more structured, easier to model with diffusion and reliably decodable into tokens. JPEG-DLM achieves the lowest mean Gen-PPL and highest throughput among recent diffusion and flow models on LM1B and OWT. At a compression rate of 0.5 on OWT, it reaches a Gen-PPL of 34.52 and approximately 2.3 times ELF's throughput. These results suggest that jointly learning compressed embeddings offers a promising path toward efficient diffusion language modeling. Code will be released soon.