diffusion pipeline parallelism

Designs, implements, and analyzes pipeline-parallel execution for diffusion models, partitioning model and generation work across devices along the generation axis and distributing denoising/time-step computations across pipeline stages. Optimizes stage scheduling, overlap, buffering and communication to minimize pipeline bubbles and to enable finer-grained pipelining between rollout and training.

diffusionpipelineparallelism

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges of deploying diffusion models at scale, including GPU memory constraints, imbalanced compute and memory loads, and the overhead and scheduling fragility introduced by staged execution across heterogeneous GPUs. To overcome these limitations, the authors propose decoupling the encoder, DiT, and decoder into independent services and organizing them into an asynchronous pipeline-parallel architecture. They further design an elastic scheduling strategy that integrates performance prediction with runtime feedback to enable efficient coordination across heterogeneous hardware. Compared to monolithic deployment, the proposed system achieves 3.4–20.5× higher throughput and up to 18.5× lower end-to-end latency, substantially improving deployment flexibility and resource efficiency.

diffusion servingdisaggregated deploymentheterogeneous GPUs

This work addresses the high computational cost of diffusion model inference and the limited scalability of existing distributed parallelization methods, which often fail to achieve linear speedup across multiple GPUs and may introduce generation artifacts. The authors propose a novel hybrid parallel framework that partitions computation based on conditional and unconditional denoising paths. By integrating conditional guidance scheduling, a hybrid data-pipeline parallelism strategy, and an adaptive parallel mode switching mechanism, the approach simultaneously optimizes inference speed and generation fidelity on both U-Net and DiT architectures. Experiments demonstrate that, using two NVIDIA RTX 3090 GPUs, the method reduces inference latency by 2.31× for SDXL and 2.07× for SD3 while preserving image quality, significantly outperforming current acceleration techniques.

conditional guidancediffusion modelsdistributed parallelism

This work addresses the lack of non-convex convergence guarantees in PipeDream-style pipeline parallelism by proposing the Randomized PipeDream (RPD) framework, for which it establishes the first rigorous non-convex convergence theory. By introducing a randomized block SGD abstraction coupled with explicit modeling of communication delays, the analysis reveals that under steady-state conditions, the delay grows quadratically with the number of pipeline stages \(S\), leading to stale gradient terms scaling as \(\Theta(S^4)\). Empirical evaluations demonstrate that RPD outperforms LocalSGD in quadratic optimization and small-scale language model training, whereas LocalSGD exhibits superior performance as \(S\) increases in logistic regression tasks, highlighting a nuanced trade-off between the two methods across different problem settings.

convergencedistributed trainingmodel parallelism

StreamDiffusion: A Pipeline-level Solution for Real-time Interactive Generation

Dec 19, 2023
AK
Akio Kodaira
🏛️ UC Berkeley | University of Tsukuba | International Christian University | Toyo University | Tokyo Institute of Technology | Tohoku University | MIT

To address the low throughput, high latency, and excessive power consumption of diffusion models in continuous interactive scenarios—such as the metaverse and real-time video streaming—this paper proposes StreamDiffusion, the first streaming diffusion generation framework designed explicitly for real-time interaction. Its core contributions are: (1) Stream Batch, a dynamic input-frame aggregation and decoupling mechanism; (2) Residual Classifier-Free Guidance (RCFG), which reduces guidance overhead while preserving generation quality; and (3) Stochastic Similarity Filtering (SSF), enabling adaptive skipping of redundant computations. The framework integrates parallel input/output queues, seamless Diffusers library compatibility, and fine-grained GPU optimizations. Evaluated on an RTX 4090, StreamDiffusion achieves 91.07 FPS, improves throughput by 59.56×, accelerates inference by 2.05×, and reduces energy consumption by 1.99–2.39×, significantly advancing the deployment of diffusion models in real-time interactive applications.

Enables real-time interactive image generation with high throughputOptimizes power consumption in continuous input scenarios like MetaverseReduces redundant computations in diffusion pipeline for faster processing

To address high inference latency and substantial inter-GPU communication overhead in generating high-resolution images with Diffusion Transformer (DiT) models, this paper proposes a block-level and layer-wise collaborative pipeline parallelism scheme. We introduce the first patch-level pipeline parallelism paradigm, integrated with a cross-diffusion-step stale feature map reuse mechanism, drastically reducing inter-GPU communication volume. The method further supports parameter sharding across devices and memory-efficient GPU memory scheduling. To our knowledge, this is the first work enabling low-latency inference of ultra-large DiT models—such as Flux.1—on an 8×L40 PCIe GPU cluster. It achieves state-of-the-art throughput and latency on PixArt-α, Stable Diffusion 3, and Flux.1, outperforming Tensor Parallelism, Sequence Parallelism, and DistriFusion by up to several orders of magnitude in communication reduction.

Enhances memory efficiency by distributing parameters across multiple GPUsOptimizes communication and computation via patch-level pipeline parallelismReduces latency in high-resolution image generation using diffusion transformers

Latest Papers

What's happening recently
View more

This work addresses the unpredictable backpressure in existing distributed DiT inference systems, which stems from neglecting the intermediate phase between communication initiation and remote visibility completion, thereby limiting effective overlap of communication and computation. To resolve this, we explicitly model this phase as a schedulable X-Stage and introduce a lightweight Burst-Gap model to characterize bursty remote memory behavior. Leveraging this insight, we redesign fused kernels to mitigate backpressure and enhance overlap efficiency. Key techniques include fine-grained device-side communication initiation, persistent GPU kernels, interleaved expert wave execution, and tile-level fusion of FlashAttention with All-to-All. Experiments across 84 configurations show that our DeepGEMM MegaMoE fused kernel achieves a 1.18× geometric mean and up to 1.62× speedup, while Ulysses attention with FlashAttention-3/4 attains peak speedups of 1.43× and 1.42×, respectively, with long-sequence steady-state performance approaching the compute-only baseline.

backpressurecommunication-computation overlapdistributed inference

This work addresses the severe communication overhead in pipeline-parallel training of large-scale diffusion models on GPU clusters, caused by long-range skip connections in the UNet architecture, which significantly limits training efficiency. To tackle this, the study introduces an automated pipeline parallelism strategy that treats skip connection locality as a primary optimization objective. It employs a skip-aware dynamic programming partitioner to colocate encoder-decoder layers on the same device and locally cache activations, thereby eliminating cross-device communication. This approach is further integrated with an integer linear programming–based bubble-efficient scheduler and a hybrid parallelism tuner to achieve end-to-end co-optimization. On communication-constrained hardware, the proposed method reduces communication volume by 89% and improves training throughput by up to 2.3× compared to state-of-the-art baselines.

communication bottleneckdiffusion modelspipeline parallelism

This work addresses the challenges of workflow scalability and low resource utilization encountered by large-scale experiments, such as those in high-energy physics, on exascale computing platforms. To overcome these limitations, we propose a multi-stage task scheduling method. By constructing a Monte Carlo simulation pipeline model alongside theoretical numerical analysis tools, this approach enables the automatic identification of optimal scheduling strategies and adaptive resource matching according to problem scale. Validation using the SBND experiment demonstrates that the proposed method effectively optimizes resource allocation for both simulation and data processing pipelines. Consequently, it significantly enhances the execution efficiency and scalability of large-scale scientific workflows deployed on exascale platforms.

Exascale computinglarge-scale experimentsresource utilization

This study addresses the absence of open standards for CPU pipeline visualization tools and the difficulty in localizing performance bottlenecks. To this end, it proposes an open-source event stream format alongside Catscan, an interactive viewer. Methodologically, this work introduces a structured event stream based on transactional relationships, integrating typed event modeling, persistent highlighting techniques, and domain-specific search algorithms to enable microarchitectural trace analysis from symptoms down to individual instructions. Furthermore, it supports resource-oriented views synchronized with comparative trace alignment. By successfully reproducing industry-grade debugging workflows, this project provides the community with production-validated microarchitectural visualization infrastructure.

CPU performance simulationmicroarchitecture debuggingopen-source tooling

Hot Scholars

AL

Alexander Long

Pluralis Research
Decentralized TrainingProtocol Learning
TA

Thalaiyasingam Ajanthan

Pluralis Research | Australian National University
Computer VisionOptimizationMachine Learning