Score
Designs and implements analyses, tooling, and policies to determine appropriate compute resource allocations (instance sizes, CPU/memory, vCPU counts, container/VM resource limits) by modeling utilization, capacity and performance tradeoffs. Produces right-sizing recommendations, thresholds for scaling or reconfiguration, and automated actions or reports to adjust resources for cost, performance, and efficiency objectives.
This study addresses inefficient virtual machine (VM) scheduling and resource allocation in enterprise cloud environments. Leveraging real-world operational telemetry from the SAP Cloud Platform—comprising 1,800 physical hosts and 48,000 VMs—we construct and publicly release, for the first time, a fine-grained, full-stack time-series telemetry dataset covering both SAP S/4HANA and general-purpose applications. Using observability tools, we collect infrastructure-level metrics across the entire stack and perform large-scale time-series analysis. Our analysis uncovers critical bottlenecks: CPU contention exceeding 40%, maximum VM ready-time latency reaching 220 seconds, CPU load imbalance affecting 99% of hosts, and sustained CPU utilization below 70% for over 80% of VMs. These findings establish the first enterprise-grade empirical foundation and reproducible dataset to guide the design of adaptive, workload-aware scheduling algorithms grounded in production realities.
This work challenges the necessity and efficacy of CPU throttling—a widely adopted resource-limiting mechanism in cloud computing. Through empirical analysis and systematic measurement in production cloud environments, the authors evaluate its impact on latency-sensitive applications, revealing that CPU limits consistently exacerbate tail latency, reduce CPU utilization, and increase per-request cost across most deployment scenarios. Methodologically, the study combines controlled microbenchmarks, production workload traces, and cross-cloud infrastructure profiling to isolate and quantify throttling-induced performance degradation. The key contributions are: (1) the first comprehensive empirical demonstration of the detrimental effects of CPU limiting; (2) a paradigm shift toward “no-limit-by-default, enable-on-demand”; (3) a reorientation of autoscaling policies from quota-based to performance-objective-driven control; and (4) empirically grounded design principles for CPU-unconstrained elastic resource management. This work critically questions long-standing industry practices and provides foundational evidence for rethinking cloud-native scheduling and pricing models.
This paper addresses the lack of systematic optimization for CPU and memory resource allocation during the Release phase of cloud-native DevOps. We propose the first pre-deployment offline performance optimization framework for microservices—distinct from mainstream auto-scaling research focused on the Ops phase. Our approach performs fine-grained resource configuration tuning *before* deployment, thereby mitigating auto-scaling failures caused by suboptimal memory provisioning. Methodologically, we integrate Bayesian optimization, statistical experimental design, and a goal-directed factor screening strategy to balance sampling cost and approximation accuracy. Extensive evaluation on the TeaStore benchmark demonstrates that our pre-deployment optimization significantly improves memory suitability and API-level resource utilization. Moreover, it empirically validates the necessity and context-dependent applicability of factor screening under varying optimization objectives.
This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.
This study addresses the challenge of enhancing productivity in supercomputing clusters and informing the design of exascale systems by analyzing job scheduling logs, GPU trace data, and domain-specific metadata from the Titan supercomputer. It systematically investigates the relationship between requested and actual resource utilization and its temporal evolution. Employing correlation analysis, clustering, and neural networks, the work presents the first comprehensive characterization of seasonal patterns in HPC resource usage and develops a transferable model for predicting resource utilization. By identifying key user behavior patterns, the research substantially improves the accuracy of forecasting future resource demands, thereby providing empirical foundations for optimizing configuration and planning of high-performance computing systems.
This work addresses the challenge of deploying multi-NUMA virtual machines, which requires aligning both virtual and physical NUMA topologies—a constraint that transforms resource allocation into a complex combinatorial optimization problem. The authors derive, for the first time, closed-form solutions for mapping symmetric 2- and 4-NUMA virtual machines onto physical servers with 4 and 8 NUMA nodes, substantially reducing scheduling complexity. By integrating NUMA topology modeling with combinatorial mathematics, the proposed approach enables efficient and accurate computation of residual capacity across a range of common configurations. This capability facilitates real-time scheduling in cloud platforms and supports large-scale resource reconfiguration with high optimization fidelity.
To address the challenge of dynamic multi-VM scheduling under hardware resource constraints in software-defined vehicles, this paper proposes a workload-aware hypervisor scenario configuration auto-generation framework. Methodologically, it introduces the first integration of domain-knowledge-guided parameter modeling with deep learning to construct a dynamic QoS prediction model; further, it designs an optimization algorithm that synthesizes chip vendors’ BSPs, heuristic rules, and system-level constraints to generate customized resource allocation schemes. The key contributions are: (1) automated and adaptive VM-level resource allocation, and (2) significant improvements in resource utilization and integration efficiency of in-vehicle virtualization systems. Experimental evaluation on real automotive platforms demonstrates a 32% reduction in development cycle time and a 27% average increase in resource utilization.
To address low resource scheduling efficiency, suboptimal hardware utilization, and high operational costs in cloud and grid environments, this paper proposes GlideinBenchmark—a novel system that extends the pilot-based architecture of GlideinWMS for automated resource discovery and fine-grained performance benchmarking. It supports user-defined metrics and enables data-driven resource matching. Built as a web application, GlideinBenchmark integrates distributed task scheduling with the HEPCloud decision engine to perform lightweight, scalable benchmarking across heterogeneous cloud and grid resources. Experimental evaluation demonstrates that our approach significantly improves resource matching accuracy, reduces average job completion time by 19.3%, and lowers computational cost by 22.7%. The system provides a reusable, intelligent scheduling infrastructure for high-performance computing environments, advancing adaptive and cost-efficient resource orchestration in hybrid infrastructures.
This paper exposes a fundamental trade-off between performance and cost in mainstream auto-scaling strategies for serverless computing: frequent instance cold starts and shutdowns incur 10–40% additional CPU overhead, while memory allocation exhibits 2–10× redundancy; existing optimizations often sacrifice significant latency. To address this, the authors develop a reproducible, transparent evaluation framework—open-sourcing a system that accurately emulates the control-plane behaviors of AWS Lambda and Google Cloud Run, integrated with real-world deployments and large-scale simulations. This enables the first systematic, quantitative characterization of latency, memory, and CPU overhead under realistic synchronous and asynchronous workloads. Key contributions include: (i) precise identification of auto-scaling efficiency bottlenecks; (ii) formulation of novel, overhead-aware scaling design principles; and (iii) provision of an empirical foundation and methodological guidance for building high-performance, cost-efficient serverless control planes.
This study addresses how cloud infrastructure failures distort performance metrics, thereby misleading autoscaling systems and leading to resource misallocation, increased costs, or degraded service reliability. Through controlled simulations, the authors systematically evaluate the impact of four common failure types—including storage and network routing issues—on both vertical and horizontal scaling strategies across varying instance configurations and SLO thresholds. The work presents the first quantitative analysis of how such failures bias scaling decisions, revealing that horizontal scaling is particularly sensitive to transient anomalies. It further proposes design principles to distinguish genuine workload changes from failure-induced artifacts. Experimental results demonstrate that storage failures can incur up to $258 in additional monthly costs under horizontal scaling, while routing anomalies consistently cause resource under-provisioning, offering empirical foundations for building fault-tolerant autoscaling mechanisms.