capacity planning

Designs and analyzes resource provisioning and scheduling configurations for service or processing systems to ensure stable operation and required throughput, including deriving maximum sustainable arrival rates and characterizing the stability region under heterogeneous resources. Builds analytical or simulation models to compute throughput capacity for given skill/resource assignments and to quantify capacity gains from collaboration, guiding provisioning, admission control, and task-assignment decisions.

capacityplanning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.65
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing scheduling theory struggles to handle multi-resource job scenarios with continuously distributed resource demands, as it relies on the assumption of finitely many job types—a simplification inconsistent with the high heterogeneity observed in real-world workloads. This work proposes the first family of throughput-optimal scheduling policies for continuous multi-resource job models, encompassing both preemptive and non-preemptive variants. The approach employs an adaptive discretization mechanism that dynamically adjusts granularity based on system load and demand distribution. By integrating throughput-optimal control, distribution-aware scheduling, and queueing optimization, the method achieves theoretical optimality while substantially improving computational efficiency. Experiments demonstrate superior performance over state-of-the-art index-based policies under both parametric distributions and real-world Google Borg traces, attaining industry-leading results.

continuous requirement distributionmultiresource-job schedulingqueueing models

This work addresses the technical and behavioral challenges of transitioning from node-exclusive to resource-aware scheduling in production-grade heterogeneous HPC systems, a shift that risks disrupting established scientific workflows. To enable seamless, non-disruptive migration, the authors propose a collaborative operational framework integrating a time-bound compatibility layer, observability-driven feedback mechanisms, and targeted user guidance. Built upon Slurm’s TRES resource model, the approach combines runtime compatibility support, job queue monitoring, and user behavior analysis to preserve workflow continuity while substantially improving scheduling efficiency. Empirical results demonstrate dramatic reductions in median queue wait times—from 277 minutes to under 3 minutes for CPU jobs and from 81 minutes to 3.4 minutes for GPU jobs—alongside high long-term adoption rates among users who embraced the new submission paradigm.

non-disruptive migrationproduction HPCresource-aware scheduling

Capacity Planning and Scheduling for Jobs with Uncertainty in Resource Usage and Duration

Jul 01, 2025
SP
Sunandita Patra
🏛️ AI Research | J.P. Morgan | CIB Athena Applied Intelligence

Financial institutions face capacity planning and job scheduling challenges in hybrid cloud and on-premise grid environments, where both resource requirements and execution durations exhibit dual uncertainty. Method: This paper proposes a co-optimization framework that jointly minimizes resource provisioning while maximizing service quality—specifically, on-time completion rate. Innovatively, it is the first to jointly model resource and duration uncertainty within capacity planning, employing a constraint programming framework based on paired sampling that integrates deterministic estimation with stochastic sampling for efficient approximate optimization. Contribution/Results: Experiments demonstrate that the method significantly reduces peak resource demand compared to manual scheduling, while maintaining a high on-time completion rate—validating its effectiveness in balancing these conflicting objectives under uncertainty.

Balance minimal resource usage and meeting job deadlinesEstimate resource needs for hybrid cloud and on-prem grid computingHandle uncertainty in job resource usage and duration

Using a market economy to provision compute resources across planet-wide clusters

May 23, 2009
MS
M. Stokely
🏛️ Google | Stanford University

To address resource supply-demand imbalances—manifesting as shortages and surpluses—across globally distributed heterogeneous computing clusters, this paper proposes a resource rationing mechanism grounded in real-world market economics. Methodologically, it introduces a periodic simulated-clock auction framework integrating utilization-driven reserve-price setting, long-term resource quota modeling, and supply-demand equilibrium pricing, enabling dynamic price signals to guide users’ autonomous job placement decisions. Its key contribution lies in being the first to systematically embed microeconomic market mechanisms into large-scale distributed resource allocation, replacing static quota or immediate-scheduling paradigms. Evaluated on the Google experimental market, the mechanism significantly incentivizes user migration toward underutilized clusters: resource utilization variance decreases by 32%, and shortage rate drops by 41%. These results empirically validate that price-based incentives can effectively drive system-level behavioral optimization and achieve global resource equilibrium.

Balancing supply-demand via simulated clock auctionsMarket-based provisioning for heterogeneous compute resourcesReducing shortages-surpluses by incentivizing resource-efficient behavior

Latest Papers

What's happening recently
View more

This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.

computing continuumconfiguration spaceheterogeneous infrastructure

This work addresses the challenge of effectively capturing and exploiting short-term scheduling flexibility while guaranteeing long-term worst-case service for multi-task streams. The authors propose a state-based scheduling framework that models worst-case service guarantees as dynamically updatable states and enforces schedulability by constraining state transitions to remain within a well-defined schedulable polytope. For the first time, they fully characterize this schedulable polytope, employ min-plus algebra for an efficient service model representation, and introduce “hyperbolic service” as a novel, dynamically extensible service form. The proposed approach significantly reduces scheduling decision complexity while strictly preserving quality-of-service guarantees, thereby enhancing practical deployability.

min-plus algebraschedulabilityscheduling flexibility

This work addresses the challenge of ensuring service continuity for multi-stage industrial workflows in B5G/6G networks, where conventional per-request QoS mechanisms fall short. To overcome this limitation, the authors propose a capability-aware collaborative planning framework that proactively exposes sustainable QoS capabilities within a finite network planning window. Industrial applications leverage this foresight to map workflow phases and submit demand trajectories, enabling workflow-level, forward-looking joint evaluation and dynamic coordination updates. By integrating network capability modeling, demand mapping, and adaptive coordination, the approach transcends traditional request-granularity constraints. Experimental validation on a real B5G system and large-scale simulations demonstrates that the proposed method significantly enhances service continuity, reduces request rejection rates, and substantially improves workflow completion rates under high network loads.

B5G/6G networkscapability-aware networkingindustrial services

This work addresses real-time task blocking and latency in low- to mid-Earth orbit distributed satellite constellations caused by stringent onboard resource constraints. To this end, the authors propose the Resource-Aware Task Allocator (RATA), a dynamic scheduling framework built upon a single-layer tree network (SLTN) architecture that jointly incorporates real-time satellite resource states, task arrival rates, and eclipse-induced energy–communication coupling constraints. The study presents the first quantitative characterization of the nonlinear relationship between constellation scale and system performance, identifying CPU availability as the primary bottleneck for task blocking and pinpointing the critical threshold at which system performance transitions from gradual degradation to abrupt collapse. Experimental results demonstrate robust energy supply under solar-aware scheduling; however, CPU bottlenecks intensify rapidly with constellation expansion, thereby establishing a practical upper bound on scalable constellation size.

Constellation ScalabilityDistributed Satellite SystemsReal-time Tasks

Hot Scholars

BH

Bowei He

City University of Hong Kong, MBZUAI
Data MiningLanguage ModelGenAI4ScienceAgentic AI
CZ

Chaojie Zhang

Microsoft
computer sciencepower managementsustainable computingjob scheduling
RG

Rohan Gandhi

Purdue University, Carnegie Mellon University, Microsoft Research
Computer Systems and NetworksSystems for LLMsAI Agents
JX

Jiarong Xing

UC Berkeley; Rice University
SystemsNetworkingSecurity