Scalable AI Inference: Performance Analysis and Optimization of AI Model Serving

📅 2026-04-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of systematic performance analysis in AI model deployment and inference, which hinders scalability and efficiency in real-world applications. Building upon BentoML, the study constructs a scalable inference system for a RoBERTa-based sentiment analysis model and identifies inference bottlenecks under three realistic traffic patterns: steady-state, bursty, and high-load scenarios. The authors propose a novel multi-level optimization framework tailored to practical deployment environments, applying coordinated improvements across runtime, service, and deployment layers. Leveraging statistical analysis, they quantify the impact of these optimizations and further evaluate the inference resilience of a single-node K3s cluster under perturbations. Experimental results demonstrate that the optimized system substantially reduces latency, increases throughput, and effectively enhances both the scalability and robustness of AI inference.

Technology Category

Machine Learning: Scalability of ML SystemsNatural Language Processing: Safety and RobustnessCognitive Modeling & Cognitive Systems: Conceptual Inference and Reasoning

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
AI research often emphasizes model design and algorithmic performance, while deployment and inference remain comparatively underexplored despite being critical for real-world use. This study addresses that gap by investigating the performance and optimization of a BentoML-based AI inference system for scalable model serving developed in collaboration with graphworks.ai. The evaluation first establishes baseline performance under three realistic workload scenarios. To ensure a fair and reproducible assessment, a pre-trained RoBERTa sentiment analysis model is used throughout the experiments. The system is subjected to traffic patterns following gamma and exponential distributions in order to emulate real-world usage conditions, including steady, bursty, and high-intensity workloads. Key performance metrics, such as latency percentiles and throughput, are collected and analyzed to identify bottlenecks in the inference pipeline. Based on the baseline results, optimization strategies are introduced at multiple levels of the serving stack to improve efficiency and scalability. The optimized system is then reevaluated under the same workload conditions, and the results are compared with the baseline using statistical analysis to quantify the impact of the applied improvements. The findings demonstrate practical strategies for achieving efficient and scalable AI inference with BentoML. The study examines how latency and throughput scale under varying workloads, how optimizations at the runtime, service, and deployment levels affect response time, and how deployment in a single-node K3s cluster influences resilience during disruptions.
Problem

Research questions and friction points this paper is trying to address.

AI inference
model serving
scalability
performance analysis
workload
Innovation

Methods, ideas, or system contributions that make the work stand out.

scalable AI inference
BentoML
model serving optimization
latency-throughput tradeoff
K3s deployment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hung Cuong Pham
Institute of Computer Science, University of Applied Sciences Ruhr West, Mülheim an der Ruhr, Germany
F
Fatih Gedikli
Institute of Computer Science, University of Applied Sciences Ruhr West, Mülheim an der Ruhr, Germany