dKV-Cache: The Cache for Diffusion Language Models

📅 2025-05-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Diffusion language models (DLMs), owing to their non-autoregressive architecture and bidirectional attention, cannot leverage conventional KV-caching for inference acceleration. Method: This paper introduces dKV-Cache—a training-agnostic, KV-cache-inspired mechanism tailored to DLMs—featuring delayed caching, conditional key-value management, and stepwise updates to align with their iterative denoising process. We propose two complementary variants: dKV-Cache-Decode (near-lossless) and dKV-Cache-Greedy (high-speed). We further identify, for the first time, insufficient context utilization during DLM inference. Contribution/Results: As a plug-and-play solution requiring no fine-tuning, dKV-Cache achieves 2–10× inference speedup across diverse tasks—including language understanding, mathematical reasoning, and code generation—while preserving model performance. It substantially narrows the practical gap between DLMs and autoregressive models.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsComputer Vision: Diffusion Models for Vision

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Diffusion Language Models (DLMs) have been seen as a promising competitor for autoregressive language models. However, diffusion language models have long been constrained by slow inference. A core challenge is that their non-autoregressive architecture and bidirectional attention preclude the key-value cache that accelerates decoding. We address this bottleneck by proposing a KV-cache-like mechanism, delayed KV-Cache, for the denoising process of DLMs. Our approach is motivated by the observation that different tokens have distinct representation dynamics throughout the diffusion process. Accordingly, we propose a delayed and conditioned caching strategy for key and value states. We design two complementary variants to cache key and value step-by-step: (1) dKV-Cache-Decode, which provides almost lossless acceleration, and even improves performance on long sequences, suggesting that existing DLMs may under-utilise contextual information during inference. (2) dKV-Cache-Greedy, which has aggressive caching with reduced lifespan, achieving higher speed-ups with quadratic time complexity at the cost of some performance degradation. dKV-Cache, in final, achieves from 2-10x speedup in inference, largely narrowing the gap between ARs and DLMs. We evaluate our dKV-Cache on several benchmarks, delivering acceleration across general language understanding, mathematical, and code-generation benchmarks. Experiments demonstrate that cache can also be used in DLMs, even in a training-free manner from current DLMs.
Problem

Research questions and friction points this paper is trying to address.

Slow inference in diffusion language models due to non-autoregressive architecture
Lack of key-value cache mechanism for bidirectional attention in DLMs
Under-utilization of contextual information during DLM inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Delayed KV-Cache for DLM denoising process
Step-by-step caching with dKV-Cache-Decode
Aggressive caching via dKV-Cache-Greedy
🔎 Similar Papers
No similar papers found.