Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

๐Ÿ“… 2026-09-22
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Flash-dLLM้€š่ฟ‡IOๆ„Ÿ็Ÿฅ็š„KV็ผ“ๅญ˜ๅ’Œๅนถ่กŒ่งฃ็ ่งฃๅ†ณๆ‰ฉๆ•ฃๅคง่ฏญ่จ€ๆจกๅž‹ๆŽจ็†ๆ•ˆ็އไฝŽ็š„้—ฎ้ข˜๏ผŒๆ้ซ˜ไบ†ๆŽจ็†้€Ÿๅบฆๅ’Œๅ†…ๅญ˜ๆ•ˆ็އใ€‚
๐Ÿ“ Abstract
Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive LLMs by enabling non-autoregressive text generation. However, their practical deployment remains limited by inefficient inference, largely due to the absence of effective Key-Value (KV) caching and scalable parallel decoding mechanisms. Existing acceleration methods typically study KV caching and parallel decoding in isolation, overlooking the I/O bottlenecks that arise when cache reuse and parallel token verification are jointly applied. In this work, we introduce $\textbf{Flash-dLLM}$, a training-free inference acceleration framework for fast and memory-efficient dLLMs. Flash-dLLM first identifies GPU memory I/O as a dominant bottleneck in KV-cache-enabled dLLM inference and addresses it with an I/O-aware fused KV-cache kernel that reduces redundant memory movement. Building on this optimized cache mechanism, Flash-dLLM further proposes an efficient KV-cache-driven draft-and-verify decoding strategy, where the dLLM itself serves as both drafter and verifier without requiring an auxiliary model. This unified design enables faster decoding while preserving generation quality and improving scalability to longer sequences and larger batch size. Extensive experiments on mathematical reasoning and code-generation benchmarks demonstrate that Flash-dLLM consistently outperforms existing state-of-the-art dLLM acceleration methods in both inference speed and memory efficiency. In particular, it achieves $5.1\times$ and $11.0\times$ speedups over prior strongest baseline Elastic-Cache on GSM8K and HumanEval, respectively.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Large Language Models
KV Caching
Parallel Decoding
Inference Efficiency
I/O Bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

IO-Aware KV Caching
Parallel Decoding
Diffusion LLMs
Inference Acceleration
Memory Efficiency
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Q
Quan Nguyen-Tri
VILA Lab, MBZUAI, Abu Dhabi, UAE
Mukul Ranjan
Mukul Ranjan
Researcher, MBZUAI
Machine Learning
Z
Zhiqiang Shen
VILA Lab, MBZUAI, Abu Dhabi, UAE