🤖 AI Summary
This study addresses the limitations of existing continuous diffusion language models, which lack variable-length generation and KV cache support while underperforming autoregressive and discrete diffusion counterparts. We propose Clock Diffusion, a framework that enables semi-autoregressive continuous diffusion through position-dependent noise scheduling. The method introduces a sliding window mechanism and a Cache Grab sampler, integrated with confidence-thresholded committing and self-speculative decoding to substantially enhance training and sampling efficiency. Experiments demonstrate that our approach achieves state-of-the-art likelihood bounds on OpenWebText. Furthermore, it significantly outperforms continuous baselines on the GSM8K benchmark and matches the performance of discrete diffusion models, effectively bridging the capability gap in continuous diffusion language modeling.
📝 Abstract
Recent works on continuous diffusion for discrete data have demonstrated performance on par with comparable discrete diffusion models. However, these continuous counterparts lack key features that are essential to practical use as language models, namely variable-length generation and support for a key-value cache, and they still lag behind the frontier of autoregressive and discrete diffusion quality. In this work, we address these limitations. We do so by introducing a model parameterization that uses position-dependent noise schedules to define semi-autoregressive (SAR) continuous diffusion language models (DLMs). Together with efficient training and sampling algorithms, we call this framework Clock Diffusion, and we present two special cases of our method: block and sliding window generation. We then define ClockDLMs, a family of Gaussian DLMs based on sliding window Clock Diffusion that attain state-of-the-art diffusion likelihood bounds on OpenWebText, even beating the performant block SAR discrete diffusion models. ClockDLMs trained on TinyGSM also substantially outperform continuous baselines on the GSM8K benchmark and match and exceed comparable SAR discrete diffusion models. Finally, building on our parameterization, we propose more efficient samplers that we dub Cache Grab, which adapt techniques from accelerated inference in discrete diffusion, such as committing tokens whose probabilities exceed a confidence threshold and self-speculative decoding, further improving our models' quality and efficiency.