Score
Designs and implements execution schedules that partition model layers into pipeline stages across hardware resources and orchestrate staged hidden‑state transfers between stages. Analyzes and optimizes these schedules to overlap inter‑device (e.g., CPU–GPU) communication with computation to hide transfer latency, preserve attention / key‑value cache locality, and minimize tail latency for long‑context requests.
本文提出一种预测引导的运行时系统,通过预测GPU池状态来优化多代理LLM工作流的物理执行图,减少延迟并提高资源利用率。
SYCL programs on multi-GPU clusters suffer from high scheduling latency and substantial critical-path overhead due to implicit memory allocation, cache-coherence operations, and dependency analysis. Method: We propose the Instruction Graph—a novel intermediate representation that fully decouples scheduling from execution. Our approach integrates speculative scheduling, adaptive virtual-buffer memory allocation, and tight integration with the Celerity runtime, enabling fully concurrent scheduling of memory management, data transfers, MPI communication, and kernel launches while moving all scheduling analysis off the critical execution path. Contribution/Results: Evaluated on a production-scale 128-GPU cluster, our method achieves excellent strong scaling, drastically reduces multi-application scheduling latency, and drives critical-path overhead nearly to zero—thereby overcoming fundamental limitations of conventional static and blocking schedulers.
Existing evaluations of pipeline parallelism scheduling strategies for large language models are limited by analytical models that neglect communication overhead and costly end-to-end experiments. This work proposes a unified evaluation framework that integrates formal modeling, tabular scheduling abstractions, and communication-aware execution simulation, enabling—for the first time—joint modeling of structured schedule representations and communication costs. Using this framework, we systematically compare GPipe, 1F1B, Chimera, and Hanayo across diverse hardware configurations, revealing that scheduling efficacy is highly dependent on the execution environment and challenging the conventional paradigm of relying solely on structural metrics such as bubble ratio. Our experiments show that GPipe and 1F1B yield similar training times, though 1F1B uses less activation memory; Chimera is advantageous only with few microbatches or highly efficient communication; and Hanayo performs well within its applicable scenarios but is sensitive to network bottlenecks.
本文提出QEffect,通过明确状态与资源契约解决低精度流水线训练中的数据流问题,保证数值更新的序列化及资源所有权,减少内存拷贝,提升训练速度。
This work addresses the significant degradation in inference throughput caused by GPU memory constraints when concurrently deploying multiple large language models on shared heterogeneous hardware, where resource scheduling, model offloading, and preemption become critical bottlenecks. Through empirical methodologies—including cross-platform performance profiling, layer-wise offloading experiments, and fine-grained decomposition of preemption overhead—the study systematically uncovers, for the first time, the nonlinear relationship between offloading and throughput decline. It further identifies model state reloading as the primary source of preemption overhead. The findings reveal that smaller models are more sensitive to reduced GPU residency, and that such overhead is jointly influenced by model architecture and hardware characteristics. These insights motivate a scheduler design that integrates model-specific sensitivity with data migration costs, offering crucial guidance for building efficient multi-model serving systems.
This study addresses the absence of open standards for CPU pipeline visualization tools and the difficulty in localizing performance bottlenecks. To this end, it proposes an open-source event stream format alongside Catscan, an interactive viewer. Methodologically, this work introduces a structured event stream based on transactional relationships, integrating typed event modeling, persistent highlighting techniques, and domain-specific search algorithms to enable microarchitectural trace analysis from symptoms down to individual instructions. Furthermore, it supports resource-oriented views synchronized with comparative trace alignment. By successfully reproducing industry-grade debugging workflows, this project provides the community with production-validated microarchitectural visualization infrastructure.
This study addresses the limited transferability of optimization knowledge in dataflow architectures caused by reliance on vendor libraries. To this end, it proposes Loom, a symbolic compiler that formulates SPMD compilation as hardware-explicit static optimization while maintaining parameters in symbolic form. By jointly optimizing schedules and parameters through CP-SAT constraint solving, legality derivation, and spatial mapping enumeration, Loom establishes the first tuning-free symbolic compilation framework, enabling cross-architecture retargetable and interpretable compiler optimizations. Evaluated on the Tenstorrent architecture, Loom matches or surpasses vendor library performance on workloads such as GEMM without requiring shape-by-shape profiling.
本文提出Para-Pipe框架,通过结合操作符并行和流水线技术优化SoC上深度学习应用的吞吐量与延迟,减少处理器间通信开销,提高能效。
This work addresses the challenges of workflow scalability and low resource utilization encountered by large-scale experiments, such as those in high-energy physics, on exascale computing platforms. To overcome these limitations, we propose a multi-stage task scheduling method. By constructing a Monte Carlo simulation pipeline model alongside theoretical numerical analysis tools, this approach enables the automatic identification of optimal scheduling strategies and adaptive resource matching according to problem scale. Validation using the SBND experiment demonstrates that the proposed method effectively optimizes resource allocation for both simulation and data processing pipelines. Consequently, it significantly enhances the execution efficiency and scalability of large-scale scientific workflows deployed on exascale platforms.
本文针对大型语言模型在完成导向任务中的调度与并行问题,提出PipeSwift方法,通过优化作业完成时间和管线并行策略,显著减少了整体作业完成时间。