🤖 AI Summary
This study addresses the low denoising efficiency and reliance on supervised objectives in diffusion language models by proposing a time-based self-supervised anchoring mechanism coupled with latent space caching. The method employs a dual-network architecture with a fusion module, utilizing a two-stage strategy to reuse semantic anchors. This preserves semantic consistency and accelerates generation without requiring supervised keyword tokens. When integrated with the DiffusionGemma-2B model, the proposed approach achieves a 49%–79% throughput improvement and a 38% reduction in computational cost on mathematical and code benchmarks, while delivering a 73% increase in empirical inference speed over the ADLM baseline.
📝 Abstract
Recent work on anchored diffusion language models improves denoising by shaping an intermediate latent space with supervised important-token targets. In this work, we introduce time-based (self-supervised) anchoring, which learns and reuses latent anchors without requiring such targets. Our key observation is that anchors encode persistent properties of the clean sequence, such as its semantic intent, global structure, or intermediate plan. Although their hidden representations become stale as the token canvas evolves, their semantic content remains useful across nearby diffusion times. This is implemented through a two-stage architecture consisting of a relatively expensive anchor network that generates the latent cache state and a lightweight denoising network that intelligently combines the cached latent state with the current state at each reverse step using a fusion module. This gives anchoring a latent-space caching interpretation: the anchor network is evaluated periodically, while its cached representation is reused across multiple reverse steps. We instantiate this framework as TADM:Post-train, which time-anchorizes pretrained DLMs, and TADM:Pretraining, which learns time-based anchors during pretraining. Applied to DiffusionGemma-26B, TADM:Post-train improves throughput by approximately 49% to 79% on several math, code, and STEM benchmarks (GSM8K, AIME26, GPQA-Diamond, LiveCodeBench-v6, HumanEval, MMLU-Pro). TADM:Pretraining reduces Transformer-layer computation by up to 38% relative to a standard single-stage DLM, achieves up to 73% higher measured throughput than ADLM.