cloud cost and scalability optimization

Designs, implements, and analyzes methods, policies, and tooling that minimize monetary cloud spend while satisfying performance, availability, and capacity requirements; this includes capacity planning, resource right‑sizing, instance/type selection, spot/commitment optimization, and cost forecasting. Builds autoscaling rules and algorithms, scheduling and provisioning strategies, monitoring and billing-analysis pipelines, and trade‑off models to automate and validate scalability and cost decisions across cloud resources and deployments.

cloudcostandscalability

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Capacity Planning and Scheduling for Jobs with Uncertainty in Resource Usage and Duration

Jul 01, 2025
SP
Sunandita Patra
🏛️ AI Research | J.P. Morgan | CIB Athena Applied Intelligence

Financial institutions face capacity planning and job scheduling challenges in hybrid cloud and on-premise grid environments, where both resource requirements and execution durations exhibit dual uncertainty. Method: This paper proposes a co-optimization framework that jointly minimizes resource provisioning while maximizing service quality—specifically, on-time completion rate. Innovatively, it is the first to jointly model resource and duration uncertainty within capacity planning, employing a constraint programming framework based on paired sampling that integrates deterministic estimation with stochastic sampling for efficient approximate optimization. Contribution/Results: Experiments demonstrate that the method significantly reduces peak resource demand compared to manual scheduling, while maintaining a high on-time completion rate—validating its effectiveness in balancing these conflicting objectives under uncertainty.

Balance minimal resource usage and meeting job deadlinesEstimate resource needs for hybrid cloud and on-prem grid computingHandle uncertainty in job resource usage and duration

The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies

Sep 03, 2025
LK
Leonid Kondrashov
🏛️ NTU | Nanjing University

This paper exposes a fundamental trade-off between performance and cost in mainstream auto-scaling strategies for serverless computing: frequent instance cold starts and shutdowns incur 10–40% additional CPU overhead, while memory allocation exhibits 2–10× redundancy; existing optimizations often sacrifice significant latency. To address this, the authors develop a reproducible, transparent evaluation framework—open-sourcing a system that accurately emulates the control-plane behaviors of AWS Lambda and Google Cloud Run, integrated with real-world deployments and large-scale simulations. This enables the first systematic, quantitative characterization of latency, memory, and CPU overhead under realistic synchronous and asynchronous workloads. Key contributions include: (i) precise identification of auto-scaling efficiency bottlenecks; (ii) formulation of novel, overhead-aware scaling design principles; and (iii) provision of an empirical foundation and methodological guidance for building high-performance, cost-efficient serverless control planes.

Analyzing serverless autoscaling performance-cost trade-offs across platformsEvaluating parameter impacts on latency and resource efficiencyQuantifying computational and memory overhead from instance churn

Existing auto-scaling frameworks suffer from prediction inaccuracy, response latency, and structural fragmentation between proactive and reactive components under highly volatile cloud workloads. To address these issues, this paper proposes OptScaler—the first unified optimization framework that synergistically integrates predictive and reactive decision-making. Its core innovation is a centralized optimization orchestrator that jointly embeds time-series forecasting, real-time self-tuning estimators, and model predictive control (MPC) with chance constraints, enabling co-optimization of resource utilization and service-level objective (SLO) compliance. This architecture overcomes the module incompatibility and insufficient robustness inherent in conventional hybrid scaling approaches. Experimental evaluations demonstrate that OptScaler reduces SLO violation rates by over 36%. Deployed at scale in Alipay’s production environment, it effectively supports elastic scaling for long-running, high-concurrency, multi-tenant applications.

Deployed at Alipay for robust payment platform supportEnhances cloud autoscaling with proactive-reactive integrationReduces SLO violations via optimized workload prediction

This paper addresses the lack of systematic optimization for CPU and memory resource allocation during the Release phase of cloud-native DevOps. We propose the first pre-deployment offline performance optimization framework for microservices—distinct from mainstream auto-scaling research focused on the Ops phase. Our approach performs fine-grained resource configuration tuning *before* deployment, thereby mitigating auto-scaling failures caused by suboptimal memory provisioning. Methodologically, we integrate Bayesian optimization, statistical experimental design, and a goal-directed factor screening strategy to balance sampling cost and approximation accuracy. Extensive evaluation on the TeaStore benchmark demonstrates that our pre-deployment optimization significantly improves memory suitability and API-level resource utilization. Moreover, it empirically validates the necessity and context-dependent applicability of factor screening under varying optimization objectives.

Address unexplored resource configuration in DevOps Release phaseCompare optimization algorithms for cost-effective near-optimal configurationsOptimize CPU and memory resource allocation for microservices pre-deployment

Latest Papers

What's happening recently
View more

This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.

computing continuumconfiguration spaceheterogeneous infrastructure

This work addresses the challenges of autoscaling in serverless computing caused by dynamic workloads, cold-start latency, and inter-function dependencies. To this end, the authors propose a dependency-aware autoscaling framework that identifies critical functions through a directed dependency graph and integrates weighted degree centrality analysis with an ensemble of lightweight multi-expert models—comprising MLP, LSTM, and CNN—combined via Bayesian heuristic probability fusion to achieve high-accuracy workload forecasting. A cost- and cold-start-aware control policy is further designed to optimize resource provisioning. Experimental results demonstrate a prediction accuracy of 99.88%, substantially outperforming existing hybrid forecasting approaches, and show significant reductions in infrastructure costs across diverse cloud pricing models while meeting performance objectives.

autoscalingcold-start latencydynamic workloads

This work addresses the inefficiencies and cost waste in cloud virtual machines caused by over-provisioning, particularly under shifting workloads that hinder dynamic tuning. The authors propose an engineer-centric, interactive instance tuning system that, for the first time, integrates zero-shot time-series forecasting models—such as Chronos-2—into non-stationary cloud environments. By leveraging large language models to generate offline recommendations, structured decision contexts, and multi-timescale suggestions, the system aligns resource allocation decisions with latency and cost constraints without requiring per-tenant model training, thereby substantially reducing operational overhead. Evaluation on seven production VMs demonstrates a 52.9% average monthly cost reduction (saving \$795) with only a 1.5% resource violation rate, achieving recommendation quality comparable to supervised baselines.

cloud instance sizingdecision alignmentoverprovisioning

This study addresses the challenges of unpredictable costs and single-region constraints associated with Spot instances in cloud services, which stem from dynamic regional pricing, variable resource availability, and interruption risks. To overcome these limitations, the authors propose an AI-driven, multi-region Spot fleet provisioning approach that integrates real-time monitoring with machine learning–based cost prediction models. Leveraging the AWS EC2 Spot Fleet API, the method enables accurate cross-region cost estimation and optimal resource allocation prior to deployment. As the first solution supporting both cross-region Spot cost forecasting and deployment optimization, this work transcends the inherent EC2 Spot restrictions of single-region operation and lack of cost predictability. Evaluated at a scale of 1,500 vCPUs, the approach achieves 99.79% cost prediction accuracy and realizes up to 64% cost savings by exploiting inter-regional price differentials.

cloud cost optimizationmulti-region provisioningprice variability

Traditional threshold-driven reactive autoscaling struggles to handle dynamic workloads, heterogeneous environments, and latency-sensitive applications, often leading to resource imbalance and performance degradation. This work proposes a predictive autoscaling framework that integrates drift awareness, uncertainty quantification, and privacy-preserving mechanisms, enabling proactive and adaptive scheduling in cloud-edge协同 environments through Kubernetes Custom Resource Definitions (CRDs) and a MAPE (Monitor-Analyze-Plan-Execute) control loop. The core contributions include a comprehensive taxonomy encompassing triggering mechanisms, target entities, prediction models, and evaluation metrics; the formulation of an Autoscaling Drift Index (ADI); and the integration of federated learning, container isolation, and feedback correction techniques. Together, these advances establish a theoretical foundation and key technical pathways for autoscaling in cloud-native and cloud-edge federated systems.

cloud-nativefederated cloud-edge computinglatency-sensitive applications