🤖 AI Summary
This work addresses the challenge that task scheduling under strong scaling is often constrained by task granularity, where scheduling overhead can dominate performance as parallelism increases, yet a systematic understanding of how algorithmic dependency structures affect scheduling efficiency remains lacking. The paper proposes a novel framework that characterizes granularity based on the dependency topology of task graphs, attributing the growth of scheduling overhead to dependency structure rather than problem size—a distinction made for the first time. Building on this insight, the authors develop a predictive model for strong-scaling limits and derive rules for selecting appropriate scheduling strategies. Through task graph analysis and overhead modeling, the approach accurately explains both gradual and abrupt scaling breakdowns observed across diverse parallel workloads, enabling informed automatic selection between static and dynamic scheduling without exhaustive empirical testing.
📝 Abstract
Task-based runtime systems provide flexible load balancing and portability for parallel scientific applications, but their strong scaling is highly sensitive to task granularity. As parallelism increases, scheduling overhead may transition from negligible to dominant, leading to rapid drops in performance for some algorithms, while remaining negligible for others. Although such effects are widely observed empirically, there is a general lack of understanding how algorithmic structure impacts whether dynamic scheduling is always beneficial. In this work, we introduce a granularity characterization framework that directly links scheduling overhead growth to task-graph dependency topology. We show that dependency structure, rather than problem size alone, governs how overhead scales with parallelism. Based on this observation, we characterize execution behavior using a simple granularity measure that indicates when scheduling overhead can be amortized by parallel computation and when scheduling overhead dominates performance. Through experimental evaluation on representative parallel workloads with diverse dependency patterns, we demonstrate that the proposed characterization explains both gradual and abrupt strong-scaling breakdowns observed in practice. We further show that overhead models derived from dependency topology accurately predict strong-scaling limits and enable a practical runtime decision rule for selecting dynamic or static execution without requiring exhaustive strong-scaling studies or extensive offline tuning.