compute-aware scheduling

Designs, builds, or analyzes scheduling algorithms and systems that jointly allocate computation and communication work while accounting for device heterogeneity and slow tasks (stragglers). This includes policies to order partitions or transmissions, overlap data transfer and computation, and co-schedule matching of compute and communication phases to minimize round duration, reduce fragmentation and overhead, and adapt to varying compute capabilities.

compute-awarescheduling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

The Merit of Simple Policies: Buying Performance With Parallelism and System Architecture

Mar 20, 2025
MY
Mert Yildiz
🏛️ University of Rome Sapienza

This paper investigates the joint optimization of server count, scheduling policy, and system architecture under a fixed computational budget to minimize average job response time. Using high-resolution traces from Google Cloud production workloads, we develop a multi-stage server cluster model and systematically compare classical policies—including Join-Idle-Queue (JIQ) and Round-Robin (RR)—against state-of-the-art size-aware schedulers. Our findings reveal: (1) an optimal critical server scale that minimizes response time; (2) in high-parallelism or multi-tier architectures, RR and JIQ significantly outperform conventional size-aware policies; and (3) parallelism degree and architectural design exert greater influence on performance than scheduling algorithm sophistication. Collectively, these results establish a new optimization paradigm wherein “architecture–parallelism” dominates over “algorithmic refinement.”

Comparing simple vs. complex dispatching policies for workload scheduling.Exploring the impact of parallelism and system architecture on performance.Optimizing job response time in cloud computing clusters.

"Two-Stagification": Job Dispatching in Large-Scale Clusters via a Two-Stage Architecture

May 05, 2025
MY
Mert Yildiz
🏛️ University of Rome Sapienza

To optimize average response time in large-scale FCFS server clusters, this paper proposes a lightweight two-stage architecture: jobs are partitioned based on an adaptive service-time threshold—short jobs are scheduled via JIQ or LWL, while long jobs employ Round-Robin (RR). This design decouples job-size sensitivity from scheduling complexity without requiring real-time server-state awareness, substantially reducing system implementation overhead. Evaluations under Weibull-synthetic workloads and Google cluster traces demonstrate that our approach significantly outperforms single-stage baselines across diverse load conditions and closely approaches the performance of state-of-the-art size-and-state-aware schedulers. The core contribution lies in empirically validating that architectural job partitioning—not scheduler complexity—is a more efficient pathway to near-optimal size-aware scheduling performance.

Improving job dispatching in large-scale clustersReducing mean response times with architectural designSeparating large and short jobs effectively

This paper investigates the impact of scheduling policies on job response time under realistic datacenter workloads. Leveraging empirical traces from Google’s production cluster, we develop a data-driven simulation framework that integrates G/G queueing theory with hierarchical job- and task-level analysis to systematically evaluate JIQ, LWL, and RR across varying cluster sizes, computational budgets, and load characteristics. Key contributions include: (1) the first quantitative identification of a performance inversion between JIQ and the size-aware policy LWL at the task level; (2) the discovery that multiple policies exhibit an optimal server count—beyond which mean response time degrades—under real-world loads; and (3) the proposal of a novel two-phase dynamic partitioning scheduler based on service-threshold criteria, which significantly reduces average response time on production traces. The study provides interpretable, reproducible theoretical foundations and practical guidance for workload-aware scheduler design.

Analyzing performance impact of cluster size and workload featuresEvaluating dispatching policies in computing clusters under real workloadsExploring task partitioning strategies to enhance scheduling performance

Workload Schedulers -- Genesis, Algorithms and Differences

Nov 13, 2025
LS
L. Sliwko
🏛️ University of Westminster

This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.

Analyzing scheduler evolution from early adoptions to modern implementationsCategorizing modern workload schedulers into three distinct classesComparing scheduling strategies across local and distributed systems

Scheduling Strategies for Partially-Replicable Task Chains on Two Types of Resources

Feb 14, 2025
DO
Diane Orhan
🏛️ University of Bordeaux | CNRS | Bordeaux INP | Inria | Sorbonne Université | LIP6

This work addresses the scheduling of partially replicable task chains (e.g., SDR communication standards) on heterogeneous multicore platforms, jointly optimizing throughput and power consumption. We formulate the problem—uniquely integrating partial replicability and big-little core co-scheduling—as a dual-resource pipelined workflow scheduling problem. To solve it, we propose: (i) FERTAC/2CATAC, a near-optimal greedy algorithm; and (ii) HeRAD, an optimal dynamic programming algorithm—both unifying pipelined and replication-based parallelism. Experiments show that FERTAC/2CATAC achieves average cycle times within <10% of HeRAD’s, with at most two additional cores overhead. On the StreamPU platform and in real-world DVB-S2 deployments, our approach attains >92% of theoretical peak throughput, significantly improving energy efficiency and scalability.

Maximizing throughput while minimizing power consumptionOptimizing task execution on big and little coresScheduling partially-replicable task chains on heterogeneous multicores

Latest Papers

What's happening recently
View more

Evaluating Rapid Makespan Predictions for Heterogeneous Systems with Programmable Logic

Oct 08, 2025
MW
Martin Wilhelm
🏛️ Otto-von-Guericke University | University of Applied Sciences

To address the challenge of rapidly evaluating the impact of dynamic task mapping adjustments on overall makespan in heterogeneous systems (CPU/GPU/FPGA), this paper proposes a lightweight prediction framework integrating abstract task graph modeling, empirical performance profiling, and analytical function fitting. The method explicitly models critical high-level factors—including inter-device data transfer overhead and hardware resource congestion—thereby bridging theoretical analysis and measured performance. Compared to conventional analytical models, our framework significantly improves cross-platform makespan prediction accuracy and is systematically validated on real heterogeneous hardware. Key contributions are: (1) the first unified execution time prediction framework supporting multiple hardware backends; (2) empirical identification of data transfer and resource congestion as dominant sources of prediction error; and (3) a scalable, empirically grounded foundation for rapid, reliable mapping decisions.

Bridging analytical predictions with real-world heterogeneous system performanceEvaluating data transfer and device congestion challenges in acceleratorsPredicting makespan impact of task mapping changes in heterogeneous systems

This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.

many-core systemsprocess-based schedulingscalability

A Real-Time Digital Twin for Adaptive Scheduling

Dec 21, 2025
YZ
Yihe Zhang
🏛️ University of Illinois Chicago | Argonne National Laboratory

HPC workloads are becoming increasingly heterogeneous, rendering traditional static heuristic schedulers inadequate for dynamic resource demands. To address this, we propose SchedTwin—the first real-time digital twin system for HPC job scheduling. It continuously ingests runtime event streams to drive high-fidelity discrete-event simulation, enabling rapid online evaluation of “what-if” scenarios across multiple scheduling policies and facilitating goal-driven, closed-loop adaptive scheduling. Deeply integrated with the PBS scheduler, SchedTwin achieves low-overhead (sub-10-second decision latency) and high-accuracy online policy optimization. Experimental evaluation in production environments demonstrates that SchedTwin significantly outperforms mainstream static schedulers—overcoming the longstanding dual bottlenecks of adaptability and timeliness inherent in conventional HPC scheduling approaches.

Adaptive scheduling for diverse HPC workloadsDynamic policy selection to meet optimization goalsReal-time digital twin guides scheduling decisions

Indirect Coflow Scheduling

Nov 16, 2025
AL
Alexander Lindermayr
🏛️ Simons Institute for the Theory of Computing | UC Berkeley | University of Pittsburgh | Arizona State University | Khoury College of Computer Sciences | Northeastern University

While large-scale data transfers in reconfigurable networks have been extensively studied, indirect cooperative flow scheduling for small-scale requests—relative to single-round transmission capacity—has long been overlooked, leading to prolonged completion times and low resource utilization. Method: This paper presents the first systematic modeling and optimization of cooperative flow scheduling under this scenario. We propose a combinatorial optimization framework integrating fractional matching and indirect routing: fractional matching enables fine-grained bandwidth allocation, while multi-hop indirect paths relax direct-connectivity constraints, supporting demand-driven elastic scheduling. Building upon theoretical schedulability analysis, we design an efficient heuristic algorithm. Results: Experiments demonstrate that, in small-scale data transfer scenarios, our approach reduces average flow completion time by 32.7% and improves link resource utilization by 41.5% over state-of-the-art methods, significantly enhancing scheduling efficiency and network adaptability for lightweight traffic.

Comparing indirect routing and fractional matchings for efficiencyDesigning algorithms optimized for small demands versus large transfersScheduling coflows for small data transfers in reconfigurable networks

This study addresses inequities and inefficiencies in computational resource allocation within federated digital research infrastructures (DRIs) by framing resource scheduling as an intersection of technical and policy considerations. Integrating policy analysis, mechanism design, and simulation experiments, the work reveals a significant disconnect between resource allocation and actual utilization. It innovatively applies Braess’s paradox to demonstrate how uncoordinated federation can exacerbate load imbalances. Leveraging algorithmic game theory, transport economics models, and simulations based on both real-world (Fresco/Anvil) and synthetic workloads—complemented by empirical investigation of international DRI allocation mechanisms—the research finds that carefully tuned, transparent scheduling heuristics can closely approximate ideal performance. While federated approaches hold promise, they require coordinated governance, and current systems notably lack effective evaluation of scientific value generated per unit of allocated resources.

algorithmic fairnessFAIR-Computefederated Digital Research Infrastructure

Hot Scholars

JT

Jacek Tabor

Profesor informatyki, Uniwersytet Jagielloński
mathematicscomputer science
ŁS

Łukasz Struski

Jagiellonian University
deep learningmachine learningclusteringlearning from missing data
MV

Marian Verhelst

Micas - ESAT - KU Leuven, Belgium
Low-energy chip designsensor fusionmachine learningcross-layer optimization
CF

Charlotte Frenkel

Assistant Professor, Delft University of Technology
Neuromorphic engineeringHardware/algorithm co-designNeuroAIOn-chip learning