Score
Designs, builds, and evaluates mechanisms, algorithms, policies, and architectures that automatically adjust compute resources (containers, nodes, or cluster capacity) in response to workload and performance metrics, including controllers, configuration, and integration points for orchestrators such as Kubernetes. Work covers autoscaling strategy and policy definition, elastic/auto-scaling infrastructure and control loops, scheduler integration, and analysis of trade-offs among performance, cost, and stability for dynamic autoscaling systems.
Autoscaling in cloud-native environments faces challenges including intricate microservice dependencies, highly dynamic and heterogeneous workloads, and poor cross-environment adaptability. This paper systematically surveys representative works published since 2020 and proposes a five-dimensional taxonomy—spanning infrastructure, architecture, scaling mechanisms, optimization objectives, and behavioral modeling—to enable fine-grained technical comparison and scenario-specific applicability analysis. We identify three emerging frontiers: large language model–driven autoscaling, microservice dependency–aware scheduling, and meta-learning–enhanced generalization—thereby bridging critical gaps in dynamic workload modeling and cross-platform adaptive scaling. By integrating performance profiling, workload feature extraction, anomaly detection, and dependency analysis, our work provides academia with a clear evolutionary roadmap and delivers to industry a practical, service-quality–aware technology selection framework that jointly optimizes resource efficiency and QoS guarantees.
Traditional threshold-driven reactive autoscaling struggles to handle dynamic workloads, heterogeneous environments, and latency-sensitive applications, often leading to resource imbalance and performance degradation. This work proposes a predictive autoscaling framework that integrates drift awareness, uncertainty quantification, and privacy-preserving mechanisms, enabling proactive and adaptive scheduling in cloud-edge协同 environments through Kubernetes Custom Resource Definitions (CRDs) and a MAPE (Monitor-Analyze-Plan-Execute) control loop. The core contributions include a comprehensive taxonomy encompassing triggering mechanisms, target entities, prediction models, and evaluation metrics; the formulation of an Autoscaling Drift Index (ADI); and the integration of federated learning, container isolation, and feedback correction techniques. Together, these advances establish a theoretical foundation and key technical pathways for autoscaling in cloud-native and cloud-edge federated systems.
Kubernetes’ native autoscaling mechanisms—relying on reactive decision-making, underutilizing application-layer signals, and employing opaque control logic—frequently violate SLOs and waste resources. To address these limitations, we propose an AIOps-driven, multi-signal collaborative autoscaling framework that jointly optimizes for SLO compliance, cost constraints, and lightweight time-series demand forecasting. This work establishes the first SLO-first, cost-aware, and safety-guaranteed unified control paradigm with inherent interpretability. By integrating multidimensional metrics into a unified model and embedding a closed-loop feedback controller, our approach ensures transparent scaling decisions and auditable operational traces. Experimental evaluation demonstrates a 31% reduction in SLO violation duration, a 24% acceleration in scaling responsiveness, and an 18% decrease in infrastructure costs—all while preserving system stability and full operational traceability.
Existing auto-scaling algorithms in edge computing suffer from poor SLA compliance, complex configuration, and delayed responsiveness. To address these challenges, this paper proposes a cloud-edge collaborative adaptive hybrid scaling mechanism. The method integrates proactive prediction with reactive feedback to enable microservice-granular, SLA-constrained dynamic resource orchestration, thereby ensuring low-latency and high-reliability service delivery. Innovatively, we design an SLA-aware elasticity policy engine that unifies a lightweight edge computing framework with hybrid cloud-edge resource coordination techniques. Experimental results demonstrate that the proposed approach significantly improves SLA attainment rate (+23.6%), reduces end-to-end latency variability (−41.2%), increases resource utilization by 18.3%, and cuts configuration parameters by 60%. Overall, it achieves an effective trade-off among performance, reliability, and operational complexity.
In microservice-based cloud-native systems, auto-scaling effectiveness is fundamentally constrained by architectural design, implementation choices, and deployment practices across the software lifecycle—factors often overlooked by existing benchmarks, leading to misleading evaluations. This paper systematically classifies and identifies critical engineering challenges affecting scaling performance according to software lifecycle phases, introducing the “lifecycle-aware scaling design” paradigm—the first of its kind. Using the Sock-Shop benchmark, we comparatively evaluate five scaling strategies: threshold-based, control-theoretic, machine-learning-driven, black-box optimization, and dependency-aware approaches. Experimental results demonstrate that holistically integrating lifecycle considerations significantly improves scaling stability (37% reduction in metric volatility) and resource efficiency (22% higher CPU utilization), whereas neglecting them causes severe performance degradation. This work bridges the gap between algorithmic auto-scaling research and real-world engineering deployment, providing a foundational methodology for production-grade scaling.
This study addresses how cloud infrastructure failures distort performance metrics, thereby misleading autoscaling systems and leading to resource misallocation, increased costs, or degraded service reliability. Through controlled simulations, the authors systematically evaluate the impact of four common failure types—including storage and network routing issues—on both vertical and horizontal scaling strategies across varying instance configurations and SLO thresholds. The work presents the first quantitative analysis of how such failures bias scaling decisions, revealing that horizontal scaling is particularly sensitive to transient anomalies. It further proposes design principles to distinguish genuine workload changes from failure-induced artifacts. Experimental results demonstrate that storage failures can incur up to $258 in additional monthly costs under horizontal scaling, while routing anomalies consistently cause resource under-provisioning, offering empirical foundations for building fault-tolerant autoscaling mechanisms.
This work addresses the inefficiencies in cloud-native autoscaling caused by misalignment between business policies and resource scheduling, as well as the lack of coordination between pod and node scaling, which often leads to resource waste and performance degradation. To resolve these issues, the authors propose MAS-H2, a three-tier hierarchical multi-agent system that introduces, for the first time, a hierarchical multi-agent architecture to cloud-native autoscaling. MAS-H2 integrates strategic, planning, and execution agents to formalize business objectives, enable joint proactive scaling of pods and nodes, and achieve zero-downtime strategy migration. Implemented as a Kubernetes Operator, the system combines time-series forecasting, global utility optimization, and multi-agent control. Experimental results on GKE demonstrate that CPU utilization remains below 40%, sustained load is reduced by over 50% compared to HPA, peak CPU load during traffic bursts drops by 55%, and zero-downtime infrastructure migration is successfully accomplished.
This work addresses the challenge of proactive scaling for containerized workloads, which is hindered by highly volatile and unpredictable resource provisioning delays. To overcome this, the authors propose ADAPT, an adaptive scaling framework that integrates an online exponentially weighted moving average (EWMA) estimator with model predictive control (MPC). Its key innovation lies in the first-time dynamic coupling of runtime cold-start delay estimation with the MPC planning horizon, enabling closed-loop self-calibration of scaling policies. The system incorporates LSTM or Prophet for workload forecasting and employs a receding optimization window. Experimental results across six representative workload types demonstrate that the MPC+LSTM configuration reduces SLA violation rates to below 5%, substantially outperforming reactive horizontal pod autoscaling (7–19%) and MPC+Prophet (up to 28.7%).
This work addresses the challenges of autoscaling in serverless computing caused by dynamic workloads, cold-start latency, and inter-function dependencies. To this end, the authors propose a dependency-aware autoscaling framework that identifies critical functions through a directed dependency graph and integrates weighted degree centrality analysis with an ensemble of lightweight multi-expert models—comprising MLP, LSTM, and CNN—combined via Bayesian heuristic probability fusion to achieve high-accuracy workload forecasting. A cost- and cold-start-aware control policy is further designed to optimize resource provisioning. Experimental results demonstrate a prediction accuracy of 99.88%, substantially outperforming existing hybrid forecasting approaches, and show significant reductions in infrastructure costs across diverse cloud pricing models while meeting performance objectives.
This work addresses significant cost inefficiencies in Kubernetes arising from resource over-provisioning and the limitations of existing autoscaling mechanisms, which often exhibit delayed responses and may obscure critical issues such as memory leaks. The authors propose a container vertical resource optimization approach grounded in a five-layer safety pipeline that introduces gating mechanisms prior to scaling decisions, with memory leak detection serving as a pivotal blocking criterion. The framework integrates SLA monitoring, circuit breakers, policy engines, and human-in-the-loop approval to ensure system reliability. Memory leaks are identified through a combination of linear regression (using R² scores) and percentile-based analysis, while scaling recommendations are generated via Holt-Winters forecasting and multi-objective Pareto optimization. The solution also incorporates conflict detection between Horizontal Pod Autoscaler (HPA) and Pod Disruption Budget (PDB) policies. Evaluated on Google Kubernetes Engine, the method achieves 20–40% cost savings, demonstrates 83% accuracy in memory leak detection, and validates reliability through 1,118 test cases covering 80.3% of the system.