Scalability Optimization in Cloud-Based AI Inference Services: Strategies for Real-Time Load Balancing and Automated Scaling

📅 2025-04-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address insufficient scalability of cloud-based AI inference services under dynamic workloads, this paper proposes a reinforcement learning (RL) and deep neural network (DNN)-coordinated adaptive optimization framework. The framework introduces a novel decentralized RL-based load distribution mechanism, integrated with DNN-driven real-time demand forecasting and closed-loop feedback control, enabling millisecond-scale load balancing and elastic resource scaling. Compared to conventional threshold- or prediction-based elasticity strategies, our approach improves load balancing efficiency by 35% and reduces end-to-end response latency by 28% under realistic workload traces. It also significantly enhances system fault tolerance and scaling responsiveness. By unifying distributed decision-making, online prediction, and adaptive control, the framework establishes a new paradigm for scalable, robust, and low-latency cloud-native AI inference services.

Technology Category

Machine Learning: Scalability of ML SystemsSearch and Optimization: Learning to SearchNatural Language Processing: Learning & Optimization for NLP

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsEconomics, Online Markets and Human Computation: Economic ramifications for generative AI infrastructure and applicationsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search engines
📝 Abstract
The rapid expansion of AI inference services in the cloud necessitates a robust scalability solution to manage dynamic workloads and maintain high performance. This study proposes a comprehensive scalability optimization framework for cloud AI inference services, focusing on real-time load balancing and autoscaling strategies. The proposed model is a hybrid approach that combines reinforcement learning for adaptive load distribution and deep neural networks for accurate demand forecasting. This multi-layered approach enables the system to anticipate workload fluctuations and proactively adjust resources, ensuring maximum resource utilisation and minimising latency. Furthermore, the incorporation of a decentralised decision-making process within the model serves to enhance fault tolerance and reduce response time in scaling operations. Experimental results demonstrate that the proposed model enhances load balancing efficiency by 35 and reduces response delay by 28, thereby exhibiting a substantial optimization effect in comparison with conventional scalability solutions.
Problem

Research questions and friction points this paper is trying to address.

Optimize scalability in cloud AI inference services
Real-time load balancing for dynamic workloads
Automated scaling to minimize latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement learning for adaptive load distribution
Deep neural networks for demand forecasting
Decentralized decision-making for fault tolerance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yihong Jin
Yihong Jin
University of Illinois at Urbana-Champaign
Machine LearningPrivacy
Z
Ze Yang
Electrical and Computer Engineering Department, University of Illinois at Urbana-Champaign