layer-wise pipeline scheduling

Designs and implements execution schedules that partition model layers into pipeline stages across hardware resources and orchestrate staged hidden‑state transfers between stages. Analyzes and optimizes these schedules to overlap inter‑device (e.g., CPU–GPU) communication with computation to hide transfer latency, preserve attention / key‑value cache locality, and minimize tail latency for long‑context requests.

layer-wisepipelinescheduling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems

Mar 13, 2025
FK
Fabian Knorr
🏛️ University of Innsbruck

SYCL programs on multi-GPU clusters suffer from high scheduling latency and substantial critical-path overhead due to implicit memory allocation, cache-coherence operations, and dependency analysis. Method: We propose the Instruction Graph—a novel intermediate representation that fully decouples scheduling from execution. Our approach integrates speculative scheduling, adaptive virtual-buffer memory allocation, and tight integration with the Celerity runtime, enabling fully concurrent scheduling of memory management, data transfers, MPI communication, and kernel launches while moving all scheduling analysis off the critical execution path. Contribution/Results: Evaluated on a production-scale 128-GPU cluster, our method achieves excellent strong scaling, drastically reduces multi-application scheduling latency, and drives critical-path overhead nearly to zero—thereby overcoming fundamental limitations of conventional static and blocking schedulers.

Enhancing memory allocation and concurrency in SYCL programs on accelerator clusters.Optimizing scheduling for high-level parallel programs on multi-GPU systems.Reducing delays in distributed-memory applications through graph-based representations.

Existing evaluations of pipeline parallelism scheduling strategies for large language models are limited by analytical models that neglect communication overhead and costly end-to-end experiments. This work proposes a unified evaluation framework that integrates formal modeling, tabular scheduling abstractions, and communication-aware execution simulation, enabling—for the first time—joint modeling of structured schedule representations and communication costs. Using this framework, we systematically compare GPipe, 1F1B, Chimera, and Hanayo across diverse hardware configurations, revealing that scheduling efficacy is highly dependent on the execution environment and challenging the conventional paradigm of relying solely on structural metrics such as bubble ratio. Our experiments show that GPipe and 1F1B yield similar training times, though 1F1B uses less activation memory; Chimera is advantageous only with few microbatches or highly efficient communication; and Hanayo performs well within its applicable scenarios but is sensitive to network bottlenecks.

communication-awareLLM trainingpipeline parallelism

This work addresses the significant degradation in inference throughput caused by GPU memory constraints when concurrently deploying multiple large language models on shared heterogeneous hardware, where resource scheduling, model offloading, and preemption become critical bottlenecks. Through empirical methodologies—including cross-platform performance profiling, layer-wise offloading experiments, and fine-grained decomposition of preemption overhead—the study systematically uncovers, for the first time, the nonlinear relationship between offloading and throughput decline. It further identifies model state reloading as the primary source of preemption overhead. The findings reveal that smaller models are more sensitive to reduced GPU residency, and that such overhead is jointly influenced by model architecture and hardware characteristics. These insights motivate a scheduler design that integrates model-specific sensitivity with data migration costs, offering crucial guidance for building efficient multi-model serving systems.

CPU-GPU offloadingGPU memory constraintsheterogeneous hardware

Latest Papers

What's happening recently
View more

This study addresses the absence of open standards for CPU pipeline visualization tools and the difficulty in localizing performance bottlenecks. To this end, it proposes an open-source event stream format alongside Catscan, an interactive viewer. Methodologically, this work introduces a structured event stream based on transactional relationships, integrating typed event modeling, persistent highlighting techniques, and domain-specific search algorithms to enable microarchitectural trace analysis from symptoms down to individual instructions. Furthermore, it supports resource-oriented views synchronized with comparative trace alignment. By successfully reproducing industry-grade debugging workflows, this project provides the community with production-validated microarchitectural visualization infrastructure.

CPU performance simulationmicroarchitecture debuggingopen-source tooling

This study addresses the limited transferability of optimization knowledge in dataflow architectures caused by reliance on vendor libraries. To this end, it proposes Loom, a symbolic compiler that formulates SPMD compilation as hardware-explicit static optimization while maintaining parameters in symbolic form. By jointly optimizing schedules and parameters through CP-SAT constraint solving, legality derivation, and spatial mapping enumeration, Loom establishes the first tuning-free symbolic compilation framework, enabling cross-architecture retargetable and interpretable compiler optimizations. Evaluated on the Tenstorrent architecture, Loom matches or surpasses vendor library performance on workloads such as GEMM without requiring shape-by-shape profiling.

auto-tuningcompiler schedulingdataflow architectures

This work addresses the challenges of workflow scalability and low resource utilization encountered by large-scale experiments, such as those in high-energy physics, on exascale computing platforms. To overcome these limitations, we propose a multi-stage task scheduling method. By constructing a Monte Carlo simulation pipeline model alongside theoretical numerical analysis tools, this approach enables the automatic identification of optimal scheduling strategies and adaptive resource matching according to problem scale. Validation using the SBND experiment demonstrates that the proposed method effectively optimizes resource allocation for both simulation and data processing pipelines. Consequently, it significantly enhances the execution efficiency and scalability of large-scale scientific workflows deployed on exascale platforms.

Exascale computinglarge-scale experimentsresource utilization

Hot Scholars

ZH

Zicong Hong

Department of Computer Science and Engineering, Hong Kong University of Science and Technology
BlockchainML SystemEdge/Cloud Computing
WL

Wayne Luk

Professor of Computer Engineering, Imperial College London
Hardware and ArchitectutreReconfigurable ComputingDesign Automation
LL

Li Li

University of Macau
Mobile ComputingCloud Computing
MX

Mingjun Xiao

University of Science and Technology of China
Mobile ComputingCrowdsensingMobile Social NetworkVechular Network
MB

Mahdi Boloursaz Mashhadi

Lecturer (Assistant Professor) at University of Surrey
Wireless CommunicationsSignal ProcessingMachine Learning