🤖 AI Summary
This work addresses the challenge of proactive scaling for containerized workloads, which is hindered by highly volatile and unpredictable resource provisioning delays. To overcome this, the authors propose ADAPT, an adaptive scaling framework that integrates an online exponentially weighted moving average (EWMA) estimator with model predictive control (MPC). Its key innovation lies in the first-time dynamic coupling of runtime cold-start delay estimation with the MPC planning horizon, enabling closed-loop self-calibration of scaling policies. The system incorporates LSTM or Prophet for workload forecasting and employs a receding optimization window. Experimental results across six representative workload types demonstrate that the MPC+LSTM configuration reduces SLA violation rates to below 5%, substantially outperforming reactive horizontal pod autoscaling (7–19%) and MPC+Prophet (up to 28.7%).
📝 Abstract
Proactive autoscaling for containerized workloads depends on knowing the provisioning delay, i.e., the time between a scaling decision and the moment new capacity is ready to serve traffic. In practice, this cold-start duration can vary substantially across environments and even across consecutive scale-out events. We present ADAPT (Adaptive Duration Approximation for Predictive Timing), an online EWMA estimator that tracks coldstart duration at runtime. ADAPT feeds a dynamic planning horizon, FH-OPT, into a Model Predictive Controller (MPC) that optimizes replica counts over a rolling window. Together, these components form a closed-loop proactive autoscaling design that adapts its lookahead based on measured provisioning delay. Evaluated across three policies (MPC+LSTM, MPC+Prophet, HPA) and six workload archetypes with five random seeds, MPC+LSTM achieves below 5% SLA violation on all workloads, compared with 7-19% for reactive HPA and up to 28.7% for MPC+Prophet on bimodal traffic.