Score
Designs, builds, or analyzes algorithms, policies, and systems that dynamically assign and reassign computational and system resources (e.g., CPU/GPU, memory, pipeline stages, and bandwidth) across tasks or streams to satisfy performance, latency/slack, and cost objectives. This encompasses online and dynamic allocation and reallocation, compute budgeting and scaling, slack-aware and playout-driven scheduling, and related strategies to preserve pipeline efficiency, resilience, and utilization as workloads change.
This study addresses the challenge of dynamic resource allocation under hard deadlines and resource constraints in multicore real-time systems by proposing a data-driven optimal control framework that explicitly accounts for timing requirements. The method pioneers the integration of real-time scheduling constraints into optimal control modeling, rigorously proves the existence of feasible solutions, and develops an efficient solving algorithm to achieve optimized resource configuration. Evaluations across four benchmarks on a physical multicore platform demonstrate that the proposed approach significantly enhances overall system performance. This work provides a solid theoretical and methodological foundation for intelligent resource management in real-time systems.
This work addresses the challenges of workflow scalability and low resource utilization encountered by large-scale experiments, such as those in high-energy physics, on exascale computing platforms. To overcome these limitations, we propose a multi-stage task scheduling method. By constructing a Monte Carlo simulation pipeline model alongside theoretical numerical analysis tools, this approach enables the automatic identification of optimal scheduling strategies and adaptive resource matching according to problem scale. Validation using the SBND experiment demonstrates that the proposed method effectively optimizes resource allocation for both simulation and data processing pipelines. Consequently, it significantly enhances the execution efficiency and scalability of large-scale scientific workflows deployed on exascale platforms.
To address the low efficiency of manual parallel workflow scheduling and poor cross-domain interoperability on heterogeneous resources within the computational continuum (IoT/edge/cloud/HPC convergence), this paper proposes the first unified workflow-driven modeling and scheduling framework tailored for the computational continuum. Our approach integrates system-and-workload co-modeling with cross-domain resource abstraction and mapping, and combines a mixed-integer linear programming (MILP) solver with lightweight heuristic algorithms. For small-scale scenarios, it achieves optimal scheduling and minimal makespan; for large-scale ones, it accelerates scheduling by 99% while maintaining solution quality within a 5–10% deviation from optimality. The framework significantly reduces end-to-end latency and improves resource utilization, thereby bridging a critical research gap in automated modeling and joint optimization for cloud–HPC collaborative scheduling.
This work addresses the challenges of training large language models across geographically distributed GPU clusters, where heterogeneous network bandwidth and regional electricity price disparities hinder existing pipeline parallelism strategies from simultaneously minimizing job completion time and power cost, while also suffering from head-of-line blocking. To overcome these limitations, the authors propose BACE-Pipe, a novel framework that jointly models real-time network utilization and job characteristics for the first time. BACE-Pipe integrates dynamic priority scheduling, bandwidth-aware path planning, and electricity-price-aware GPU allocation to co-optimize pipeline execution across regions. Experimental results demonstrate that the approach effectively mitigates head-of-line blocking and significantly reduces both latency and operational cost in multi-tenant environments, achieving 27.9%–64.7% lower average job completion time and 12.6%–30.6% reduction in total electricity expenditure.
HPC workloads are becoming increasingly heterogeneous, rendering traditional static heuristic schedulers inadequate for dynamic resource demands. To address this, we propose SchedTwin—the first real-time digital twin system for HPC job scheduling. It continuously ingests runtime event streams to drive high-fidelity discrete-event simulation, enabling rapid online evaluation of “what-if” scenarios across multiple scheduling policies and facilitating goal-driven, closed-loop adaptive scheduling. Deeply integrated with the PBS scheduler, SchedTwin achieves low-overhead (sub-10-second decision latency) and high-accuracy online policy optimization. Experimental evaluation in production environments demonstrates that SchedTwin significantly outperforms mainstream static schedulers—overcoming the longstanding dual bottlenecks of adaptability and timeliness inherent in conventional HPC scheduling approaches.
This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.
This study addresses the limitations of static cache replacement and prefetching policies in conventional processors, which struggle to maintain optimal performance across diverse execution phases. For the first time, it systematically evaluates the potential of dynamic policy selection by analyzing 490 execution phases from 49 benchmark programs using the ChampSim simulator. The results demonstrate that static policies incur an average IPC loss of 1.54%, whereas dynamically switching between two carefully selected policies reduces this loss to just 0.11%. Moreover, such dynamic adaptation achieves near-ideal performance in 52.65% of the phases, closely approaching the theoretical upper bound. This work validates the efficacy of dynamic policy switching and establishes a new paradigm for enhancing single-threaded performance.
This work addresses the challenge of maximizing end-to-end success probability in structured agent workflows under hard constraints on budget and deadline. The authors propose Monte Carlo Combinatorial Planning (MCPP), a lightweight closed-loop planner that dynamically replans during execution in response to observations. MCPP employs a finite-horizon stochastic online allocation model with parallel sampling and leverages Monte Carlo simulation to estimate, in real time, the probability of successful task completion under the given constraints. Experimental results demonstrate that MCPP significantly outperforms strong baseline methods on the CodeFlow and ProofFlow benchmarks, consistently achieving higher task completion rates across diverse budget–deadline configurations. These findings validate MCPP’s effectiveness and robustness in resource-constrained scenarios.
This work addresses the limitations of traditional numerical array programs, which rely on manual parallelization constrained by static optimizations or explicit annotations, resulting in coarse-grained parallelism and poor adaptability to heterogeneous hardware. The paper proposes a self-optimizing Virtual Processor (VP) that automatically and dynamically parallelizes entire program regions at runtime through a decentralized network of collaborating execution segments, without developer intervention. Its key innovation lies in parallelizing and distributing the scheduling process itself, integrating dependency-driven local decisions, heterogeneity-aware task placement and data movement, and support from the ILNumerics.ONAL instruction set. This approach preserves sequential semantics while enabling automatic parallelism extraction across large-scale program regions, achieving low-latency strong scaling on local heterogeneous systems for a broad range of workloads—from latency-sensitive small operations to large data-parallel tasks—without requiring explicit parallel programming.
This study addresses the trade-off between energy efficiency and latency for AI/ML workloads in multi-instance GPU (MIG) environments by proposing a dynamic repartitioning scheduling framework tailored to individual MIG instances. The framework integrates scheduler selection with a reinforcement learning–based dynamic repartitioning strategy, marking the first application of reinforcement learning to MIG reconfiguration decisions. It uncovers optimal GPU partitioning patterns under varying temporal and queue-state conditions, enabling predictive and automatic adjustments. Experimental evaluation using real-world diurnal workload traces from data centers demonstrates that the proposed approach improves the combined metric of energy consumption and task latency by 26%, 31%, and 68% compared to twice-daily repartitioning, static partitioning, and no partitioning schemes, respectively.