adaptive resource allocation

Designs, builds, or analyzes algorithms, policies, and systems that dynamically assign and reassign computational and system resources (e.g., CPU/GPU, memory, pipeline stages, and bandwidth) across tasks or streams to satisfy performance, latency/slack, and cost objectives. This encompasses online and dynamic allocation and reallocation, compute budgeting and scaling, slack-aware and playout-driven scheduling, and related strategies to preserve pipeline efficiency, resilience, and utilization as workloads change.

adaptiveresourceallocation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$209K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of dynamic resource allocation under hard deadlines and resource constraints in multicore real-time systems by proposing a data-driven optimal control framework that explicitly accounts for timing requirements. The method pioneers the integration of real-time scheduling constraints into optimal control modeling, rigorously proves the existence of feasible solutions, and develops an efficient solving algorithm to achieve optimized resource configuration. Evaluations across four benchmarks on a physical multicore platform demonstrate that the proposed approach significantly enhances overall system performance. This work provides a solid theoretical and methodological foundation for intelligent resource management in real-time systems.

Dynamic Resource AllocationHard Deadline ConstraintsMulticore Real-time Systems

This work addresses the challenges of workflow scalability and low resource utilization encountered by large-scale experiments, such as those in high-energy physics, on exascale computing platforms. To overcome these limitations, we propose a multi-stage task scheduling method. By constructing a Monte Carlo simulation pipeline model alongside theoretical numerical analysis tools, this approach enables the automatic identification of optimal scheduling strategies and adaptive resource matching according to problem scale. Validation using the SBND experiment demonstrates that the proposed method effectively optimizes resource allocation for both simulation and data processing pipelines. Consequently, it significantly enhances the execution efficiency and scalability of large-scale scientific workflows deployed on exascale platforms.

Exascale computinglarge-scale experimentsresource utilization

Workflow-Driven Modeling for the Compute Continuum: An Optimization Approach to Automated System and Workload Scheduling

May 18, 2025
AK
Aasish Kumar Sharma
🏛️ Georg-August-Universität Göttingen | GWDG | Zuse Institute

To address the low efficiency of manual parallel workflow scheduling and poor cross-domain interoperability on heterogeneous resources within the computational continuum (IoT/edge/cloud/HPC convergence), this paper proposes the first unified workflow-driven modeling and scheduling framework tailored for the computational continuum. Our approach integrates system-and-workload co-modeling with cross-domain resource abstraction and mapping, and combines a mixed-integer linear programming (MILP) solver with lightweight heuristic algorithms. For small-scale scenarios, it achieves optimal scheduling and minimal makespan; for large-scale ones, it accelerates scheduling by 99% while maintaining solution quality within a 5–10% deviation from optimality. The framework significantly reduces end-to-end latency and improves resource utilization, thereby bridging a critical research gap in automated modeling and joint optimization for cloud–HPC collaborative scheduling.

Enhancing workload efficiency across heterogeneous compute resourcesOptimizing automated scheduling in IoT-Edge-Cloud-HPC continuumReducing latency and overhead in cloud-HPC resource integration

This work addresses the challenges of training large language models across geographically distributed GPU clusters, where heterogeneous network bandwidth and regional electricity price disparities hinder existing pipeline parallelism strategies from simultaneously minimizing job completion time and power cost, while also suffering from head-of-line blocking. To overcome these limitations, the authors propose BACE-Pipe, a novel framework that jointly models real-time network utilization and job characteristics for the first time. BACE-Pipe integrates dynamic priority scheduling, bandwidth-aware path planning, and electricity-price-aware GPU allocation to co-optimize pipeline execution across regions. Experimental results demonstrate that the approach effectively mitigates head-of-line blocking and significantly reduces both latency and operational cost in multi-tenant environments, achieving 27.9%–64.7% lower average job completion time and 12.6%–30.6% reduction in total electricity expenditure.

Bandwidth HeterogeneityElectricity CostGeo-Distributed Training

A Real-Time Digital Twin for Adaptive Scheduling

Dec 21, 2025
YZ
Yihe Zhang
🏛️ University of Illinois Chicago | Argonne National Laboratory

HPC workloads are becoming increasingly heterogeneous, rendering traditional static heuristic schedulers inadequate for dynamic resource demands. To address this, we propose SchedTwin—the first real-time digital twin system for HPC job scheduling. It continuously ingests runtime event streams to drive high-fidelity discrete-event simulation, enabling rapid online evaluation of “what-if” scenarios across multiple scheduling policies and facilitating goal-driven, closed-loop adaptive scheduling. Deeply integrated with the PBS scheduler, SchedTwin achieves low-overhead (sub-10-second decision latency) and high-accuracy online policy optimization. Experimental evaluation in production environments demonstrates that SchedTwin significantly outperforms mainstream static schedulers—overcoming the longstanding dual bottlenecks of adaptability and timeliness inherent in conventional HPC scheduling approaches.

Adaptive scheduling for diverse HPC workloadsDynamic policy selection to meet optimization goalsReal-time digital twin guides scheduling decisions

Latest Papers

What's happening recently
View more

This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.

many-core systemsprocess-based schedulingscalability

This study addresses the limitations of static cache replacement and prefetching policies in conventional processors, which struggle to maintain optimal performance across diverse execution phases. For the first time, it systematically evaluates the potential of dynamic policy selection by analyzing 490 execution phases from 49 benchmark programs using the ChampSim simulator. The results demonstrate that static policies incur an average IPC loss of 1.54%, whereas dynamically switching between two carefully selected policies reduces this loss to just 0.11%. Moreover, such dynamic adaptation achieves near-ideal performance in 52.65% of the phases, closely approaching the theoretical upper bound. This work validates the efficacy of dynamic policy switching and establishes a new paradigm for enhancing single-threaded performance.

cache replacementdynamic policy selectionout-of-order pipeline

This work addresses the challenge of maximizing end-to-end success probability in structured agent workflows under hard constraints on budget and deadline. The authors propose Monte Carlo Combinatorial Planning (MCPP), a lightweight closed-loop planner that dynamically replans during execution in response to observations. MCPP employs a finite-horizon stochastic online allocation model with parallel sampling and leverages Monte Carlo simulation to estimate, in real time, the probability of successful task completion under the given constraints. Experimental results demonstrate that MCPP significantly outperforms strong baseline methods on the CodeFlow and ProofFlow benchmarks, consistently achieving higher task completion rates across diverse budget–deadline configurations. These findings validate MCPP’s effectiveness and robustness in resource-constrained scenarios.

agentic workflowsbudget constraintconstraint-driven

This work addresses the limitations of traditional numerical array programs, which rely on manual parallelization constrained by static optimizations or explicit annotations, resulting in coarse-grained parallelism and poor adaptability to heterogeneous hardware. The paper proposes a self-optimizing Virtual Processor (VP) that automatically and dynamically parallelizes entire program regions at runtime through a decentralized network of collaborating execution segments, without developer intervention. Its key innovation lies in parallelizing and distributing the scheduling process itself, integrating dependency-driven local decisions, heterogeneity-aware task placement and data movement, and support from the ILNumerics.ONAL instruction set. This approach preserves sequential semantics while enabling automatic parallelism extraction across large-scale program regions, achieving low-latency strong scaling on local heterogeneous systems for a broad range of workloads—from latency-sensitive small operations to large data-parallel tasks—without requiring explicit parallel programming.

automatic optimizationheterogeneous hardwarenumerical array programs

This study addresses the trade-off between energy efficiency and latency for AI/ML workloads in multi-instance GPU (MIG) environments by proposing a dynamic repartitioning scheduling framework tailored to individual MIG instances. The framework integrates scheduler selection with a reinforcement learning–based dynamic repartitioning strategy, marking the first application of reinforcement learning to MIG reconfiguration decisions. It uncovers optimal GPU partitioning patterns under varying temporal and queue-state conditions, enabling predictive and automatic adjustments. Experimental evaluation using real-world diurnal workload traces from data centers demonstrates that the proposed approach improves the combined metric of energy consumption and task latency by 26%, 31%, and 68% compared to twice-daily repartitioning, static partitioning, and no partitioning schemes, respectively.

AI/ML WorkloadsDynamic RepartitioningEnergy-Efficient Scheduling

Hot Scholars

FF

Fangcheng Fu

Shanghai Jiao Tong University
machine learningdeep learningMLSysdistributed computation
UF

Uriel Feige

Professor of Computer Science, Weizmann Institute of Science
HX

Hongli Xu

University of Science and Technology of China
Software Defined NetworkCooperative CommunicationSensor Networks
WN

Wei Ni

FIEEE, AAIA Fellow, Senior Principal Scientist & Conjoint Professor, CSIRO/UNSW
6G security and privacyconnected and trusted intelligenceapplied AI/ML
KY

Kejiang Ye

Professor, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Cloud ComputingAI SystemsIndustrial Internet