SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models

πŸ“… 2026-10-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the computational redundancy inherent in multi-branch speculative decoding for diffusion large language models by proposing an algorithm-system co-optimization framework. Methodologically, it exploits hidden-state similarity across branches, eliminating spatial redundancy through folded attention and feed-forward network (FFN) reuse from parent branchesβ€”an approach orthogonal to temporal caching techniques. Furthermore, token-level residual gating and custom Triton sparse execution kernels are designed to enable efficient verification. Experimental results demonstrate that the proposed method achieves lossless generation quality while improving throughput by 1.64Γ— over Spiffy and 1.99Γ— compared to vanilla decoding.
πŸ“ Abstract
Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy across denoising steps, we identify a complementary redundancy axis within each speculative verification step: multi-branch computational redundancy. During speculative verification, draft branches inherit most tokens from their parents while unmasking a small set of additional positions, causing large portions of hidden states to remain highly similar across branches. We propose SpecFold, an algorithm-system co-design that exploits this multi-branch redundancy to reduce the cost of multi-branch speculative verification. Algorithmically, SpecFold performs token-level residual gating and selectively reuses parent computation through folded attention and FFN while preserving residual hidden states. Systemically, a Triton kernel implementation translates this fine-grained reuse into end-to-end throughput gains through efficient sparse multi-branch execution. SpecFold is orthogonal to temporal caching and compatible with existing DLLM speculation strategies. Across two DLLM families, five models, and five standard benchmarks, SpecFold achieves up to 1.64x throughput over Spiffy and up to 1.99x over vanilla decoding, while maintaining comparable task performance.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Large Language Models
Speculative Decoding
Multi-Branch Redundancy
Computational Efficiency
Throughput Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Diffusion Language Models
Multi-Branch Redundancy
Algorithm-System Co-design
Triton Kernel
πŸ”Ž Similar Papers
No similar papers found.