Score
Designs, builds, or analyzes scheduling algorithms and systems that jointly allocate computation and communication work while accounting for device heterogeneity and slow tasks (stragglers). This includes policies to order partitions or transmissions, overlap data transfer and computation, and co-schedule matching of compute and communication phases to minimize round duration, reduce fragmentation and overhead, and adapt to varying compute capabilities.
This paper investigates the joint optimization of server count, scheduling policy, and system architecture under a fixed computational budget to minimize average job response time. Using high-resolution traces from Google Cloud production workloads, we develop a multi-stage server cluster model and systematically compare classical policies—including Join-Idle-Queue (JIQ) and Round-Robin (RR)—against state-of-the-art size-aware schedulers. Our findings reveal: (1) an optimal critical server scale that minimizes response time; (2) in high-parallelism or multi-tier architectures, RR and JIQ significantly outperform conventional size-aware policies; and (3) parallelism degree and architectural design exert greater influence on performance than scheduling algorithm sophistication. Collectively, these results establish a new optimization paradigm wherein “architecture–parallelism” dominates over “algorithmic refinement.”
To optimize average response time in large-scale FCFS server clusters, this paper proposes a lightweight two-stage architecture: jobs are partitioned based on an adaptive service-time threshold—short jobs are scheduled via JIQ or LWL, while long jobs employ Round-Robin (RR). This design decouples job-size sensitivity from scheduling complexity without requiring real-time server-state awareness, substantially reducing system implementation overhead. Evaluations under Weibull-synthetic workloads and Google cluster traces demonstrate that our approach significantly outperforms single-stage baselines across diverse load conditions and closely approaches the performance of state-of-the-art size-and-state-aware schedulers. The core contribution lies in empirically validating that architectural job partitioning—not scheduler complexity—is a more efficient pathway to near-optimal size-aware scheduling performance.
This paper investigates the impact of scheduling policies on job response time under realistic datacenter workloads. Leveraging empirical traces from Google’s production cluster, we develop a data-driven simulation framework that integrates G/G queueing theory with hierarchical job- and task-level analysis to systematically evaluate JIQ, LWL, and RR across varying cluster sizes, computational budgets, and load characteristics. Key contributions include: (1) the first quantitative identification of a performance inversion between JIQ and the size-aware policy LWL at the task level; (2) the discovery that multiple policies exhibit an optimal server count—beyond which mean response time degrades—under real-world loads; and (3) the proposal of a novel two-phase dynamic partitioning scheduler based on service-threshold criteria, which significantly reduces average response time on production traces. The study provides interpretable, reproducible theoretical foundations and practical guidance for workload-aware scheduler design.
This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.
This work addresses the scheduling of partially replicable task chains (e.g., SDR communication standards) on heterogeneous multicore platforms, jointly optimizing throughput and power consumption. We formulate the problem—uniquely integrating partial replicability and big-little core co-scheduling—as a dual-resource pipelined workflow scheduling problem. To solve it, we propose: (i) FERTAC/2CATAC, a near-optimal greedy algorithm; and (ii) HeRAD, an optimal dynamic programming algorithm—both unifying pipelined and replication-based parallelism. Experiments show that FERTAC/2CATAC achieves average cycle times within <10% of HeRAD’s, with at most two additional cores overhead. On the StreamPU platform and in real-world DVB-S2 deployments, our approach attains >92% of theoretical peak throughput, significantly improving energy efficiency and scalability.
To address the challenge of rapidly evaluating the impact of dynamic task mapping adjustments on overall makespan in heterogeneous systems (CPU/GPU/FPGA), this paper proposes a lightweight prediction framework integrating abstract task graph modeling, empirical performance profiling, and analytical function fitting. The method explicitly models critical high-level factors—including inter-device data transfer overhead and hardware resource congestion—thereby bridging theoretical analysis and measured performance. Compared to conventional analytical models, our framework significantly improves cross-platform makespan prediction accuracy and is systematically validated on real heterogeneous hardware. Key contributions are: (1) the first unified execution time prediction framework supporting multiple hardware backends; (2) empirical identification of data transfer and resource congestion as dominant sources of prediction error; and (3) a scalable, empirically grounded foundation for rapid, reliable mapping decisions.
This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.
HPC workloads are becoming increasingly heterogeneous, rendering traditional static heuristic schedulers inadequate for dynamic resource demands. To address this, we propose SchedTwin—the first real-time digital twin system for HPC job scheduling. It continuously ingests runtime event streams to drive high-fidelity discrete-event simulation, enabling rapid online evaluation of “what-if” scenarios across multiple scheduling policies and facilitating goal-driven, closed-loop adaptive scheduling. Deeply integrated with the PBS scheduler, SchedTwin achieves low-overhead (sub-10-second decision latency) and high-accuracy online policy optimization. Experimental evaluation in production environments demonstrates that SchedTwin significantly outperforms mainstream static schedulers—overcoming the longstanding dual bottlenecks of adaptability and timeliness inherent in conventional HPC scheduling approaches.
While large-scale data transfers in reconfigurable networks have been extensively studied, indirect cooperative flow scheduling for small-scale requests—relative to single-round transmission capacity—has long been overlooked, leading to prolonged completion times and low resource utilization. Method: This paper presents the first systematic modeling and optimization of cooperative flow scheduling under this scenario. We propose a combinatorial optimization framework integrating fractional matching and indirect routing: fractional matching enables fine-grained bandwidth allocation, while multi-hop indirect paths relax direct-connectivity constraints, supporting demand-driven elastic scheduling. Building upon theoretical schedulability analysis, we design an efficient heuristic algorithm. Results: Experiments demonstrate that, in small-scale data transfer scenarios, our approach reduces average flow completion time by 32.7% and improves link resource utilization by 41.5% over state-of-the-art methods, significantly enhancing scheduling efficiency and network adaptability for lightweight traffic.
This study addresses inequities and inefficiencies in computational resource allocation within federated digital research infrastructures (DRIs) by framing resource scheduling as an intersection of technical and policy considerations. Integrating policy analysis, mechanism design, and simulation experiments, the work reveals a significant disconnect between resource allocation and actual utilization. It innovatively applies Braess’s paradox to demonstrate how uncoordinated federation can exacerbate load imbalances. Leveraging algorithmic game theory, transport economics models, and simulations based on both real-world (Fresco/Anvil) and synthetic workloads—complemented by empirical investigation of international DRI allocation mechanisms—the research finds that carefully tuned, transparent scheduling heuristics can closely approximate ideal performance. While federated approaches hold promise, they require coordinated governance, and current systems notably lack effective evaluation of scientific value generated per unit of allocated resources.