ParaAnya: Accelerating Parallel Diffusion Sampling with Plug-and-Play Output Caching

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational redundancy and efficiency bottlenecks caused by the sliding window mechanism in parallel diffusion sampling. We propose a plug-and-play output caching strategy that constructs input-output pair caches to reuse overlapping timestep results, dispatching only cache-missed tasks to the GPU and thereby eliminating redundant neural function evaluations (NFEs). This algorithm-agnostic caching mechanism significantly reduces computational overhead while preserving the structural integrity of the underlying update rules. Experiments based on Stable Diffusion v1.5 demonstrate that our method achieves up to 2.43× parallel acceleration with a 70.1% reduction in NFEs. Furthermore, in single-GPU scenarios, it yields a 5.62× speedup while maintaining comparable CLIP scores.
📝 Abstract
Diffusion models have achieved remarkable success in generative tasks, but their inherently sequential sampling process introduces a severe computational bottleneck. Recent Parallel-in-Time (PinT) solvers attempt to mitigate this by parallelizing generation across a sliding window of timesteps, advancing the window only when step-wise changes stabilize. However, this overlapping window mechanism forces the network to repeatedly evaluate the same timesteps. When the input variations between iterations are minimal, these redundant evaluations lead to significant computational waste. To address this inefficiency, we propose ParaAnya, an output cache mechanism agnostic to the parallel sampling algorithm that can reduce the number of function evaluations (NFE). ParaAnya caches input-output pairs of diffusion models and reuses the cached output at overlapping timesteps. By dispatching only cache-miss timesteps to GPU workers, our approach eliminates redundant computation while preserving the structure of the underlying algorithms' update rules. We integrate ParaAnya into four representative parallel sampling algorithms and evaluate its performance on Stable Diffusion v1.5. Across four parallel samplers evaluated with DDIM on eight GPUs, ParaAnya provides $1.30$--$2.43\times$ speedups over their uncached counterparts and reduces NFE by up to 70.1\%, reaching up to a $5.62\times$ speedup over single-GPU serial sampling while maintaining comparable CLIP scores.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Models
Parallel Sampling
Computational Bottleneck
Redundant Computation
Function Evaluations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parallel Diffusion Sampling
Output Caching
Plug-and-Play
Function Evaluations Reduction
Parallel-in-Time Solvers
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chee-En Yu
National Taiwan University, Taiwan
X
Xiao-Xi Tan
National Taiwan University, Taiwan
Yi-Cheng Lin
Yi-Cheng Lin
National Taiwan University
Speech ProcessingMachine LearningFairness
Y
Yun-Shao Tsai
National Taiwan University, Taiwan
C
Chee-An Yu
University of Southern California, USA
Hung-yi Lee
Hung-yi Lee
National Taiwan University
deep learningspoken language understandingspeech processing