model serving

Designs and implements infrastructure and runtime components that expose trained models as reliable, low‑latency, and scalable services for inference. This includes building APIs/endpoints, request routing and batching, caching, versioning and rollback, autoscaling and load balancing, deployment automation, and monitoring/observability to meet latency, throughput, and availability requirements.

modelserving

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.77
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge that existing AI systems struggle to dynamically observe, intervene in, and optimize agent behavior at runtime, making it difficult to simultaneously achieve high task success rates, low latency, token efficiency, reliability, and safety. To overcome this limitation, the paper proposes a novel runtime infrastructure layer situated between the model and the application, which treats the AI execution process itself as an optimizable object—departing from conventional approaches that restrict optimization to static model or log-level adjustments. This layer enables proactive intervention and multi-dimensional performance co-optimization through mechanisms such as runtime monitoring, real-time inference, adaptive memory management, fault recovery, and policy enforcement. Experimental results demonstrate that the proposed approach significantly enhances the holistic performance of long-horizon agent workflows across task success rate, response latency, token efficiency, system reliability, and safety compliance.

Agent ExecutionAI RuntimeLong-horizon Workflows

Scalable Engine and the Performance of Different LLM Models in a SLURM based HPC architecture

Aug 25, 2025
AD
Anderson de Lima Luiz
🏛️ AImotion Bavaria | Technische Hochschule Ingolstadt

To address challenges in high-performance computing (HPC) environments—including heterogeneous LLM deployment, inflexible resource scheduling, and significant performance volatility under multi-model concurrent inference—this paper proposes a scalable LLM inference engine architecture built atop SLURM. The architecture integrates containerized microservices with dynamic resource orchestration, enabling fine-grained, coordinated allocation of CPU, GPU, and memory resources, and provides unified access via RESTful APIs to support both batch and interactive inference workloads. A novel multi-step “tribunal” refinement workflow is introduced to enhance fault tolerance and operational flexibility. Experiments on Llama-series models across multi-node HPC clusters demonstrate sub-50 ms latency and 128 concurrent requests for smaller models (e.g., Llama-3-8B), and stable dual-concurrent execution for large models (e.g., Llama-3-70B), with low scheduling overhead and strong horizontal scalability. The system has been successfully deployed in production applications, including retrieval-augmented generation chatbots.

Developing scalable HPC architecture for efficient LLM deploymentEvaluating performance metrics across various model sizesOptimizing resource allocation for heterogeneous models in clusters

Fast and Live Model Auto Scaling with O(1) Host Caching

Dec 23, 2024
DZ
Dingyan Zhang
🏛️ Shanghai Jiao Tong University | Huawei Cloud

In large language model (LLM) serving, autoscaling faces a fundamental trade-off between high scaling latency and substantial parameter caching overhead: conventional approaches rely on local parameter caches, causing service interruption during model loading and suffering from data-plane bottlenecks in cross-host scaling. This paper proposes a local-cache-free, millisecond-scale real-time autoscaling framework. It introduces the first O(1)-complexity network-based direct parameter transmission mechanism, enabling zero-copy parameter loading over GPU-to-GPU high-speed interconnects. We design a layer-granularity dynamic collaborative execution architecture supporting fine-grained load migration and multicast-based parameter distribution. Integrated cooperative inference scheduling eliminates cold-start delays. Experiments show up to 86% reduction in tail latency, achieving performance close to the ideal full-host, full-parameter caching configuration—while completely eliminating local parameter storage overhead.

Enable live scaling via fine-grained layer-level inferenceOptimize model autoscaling speed with minimal cachingReduce GPU usage and latency in serving systems

Current large language model inference systems lack the capability to dynamically adjust model parallelism topologies at runtime, necessitating service restarts under varying workloads and causing multi-minute disruptions, loss of KV cache, and substantial recomputation overhead. This work proposes ReMP, the first framework enabling online elastic reconfiguration of combined tensor and pipeline parallelism. By decoupling topology from execution state, designing a two-dimensional KV cache migration mechanism, and orchestrating an end-to-end reconfiguration pipeline, ReMP reduces topology switching latency to 1–7 seconds across 7B–70B models—orders of magnitude faster than restarting. This significantly improves time-to-first-token (TTFT), time per output token (TPOT), and throughput under dynamic workloads.

DowntimeKV CacheLLM Serving

Latest Papers

What's happening recently
View more

This study addresses the lack of systematic empirical investigation into the deployment of large language model (LLM) inference frameworks in real-world software systems. Through large-scale analysis of open-source projects, combining static code analysis with metadata mining, this work systematically characterizes the adoption patterns and integration strategies of prominent frameworks—including vLLM, SGLang, TensorRT-LLM, LMDeploy, and FlashInfer—and examines their relationships with model type, scale, modality, and deployment environment. The findings reveal that vLLM exhibits the highest adoption rate, while multi-framework co-deployment remains limited yet holds complementary potential. Framework selection is strongly driven by model characteristics and application scenarios, effectively supporting diverse system designs such as reinforcement learning inference, multimodal generation, and microservice architectures.

empirical studyframework adoptionLLM serving

This work addresses the lack of systematic performance analysis in AI model deployment and inference, which hinders scalability and efficiency in real-world applications. Building upon BentoML, the study constructs a scalable inference system for a RoBERTa-based sentiment analysis model and identifies inference bottlenecks under three realistic traffic patterns: steady-state, bursty, and high-load scenarios. The authors propose a novel multi-level optimization framework tailored to practical deployment environments, applying coordinated improvements across runtime, service, and deployment layers. Leveraging statistical analysis, they quantify the impact of these optimizations and further evaluate the inference resilience of a single-node K3s cluster under perturbations. Experimental results demonstrate that the optimized system substantially reduces latency, increases throughput, and effectively enhances both the scalability and robustness of AI inference.

AI inferencemodel servingperformance analysis

Existing large model training data pipelines struggle to simultaneously ensure batch semantics, fault isolation, and consistency. This work proposes an agent-free, object-storage-native training data plane featuring three core innovations: a Transactional Global Batch (TGB) abstraction that guarantees training consistency, a storage-layer-embedded garbage collection mechanism aligning producer states with distributed checkpoints, and a communication-free Decentralized Adaptive Commit (DAC) algorithm. Leveraging lakehouse ACID semantics in object storage and distributed checkpointing, the system achieves higher throughput than colocated loaders and Kafka, lower read latency, and full fault isolation across 64-GPU multimodal pretraining and supervised fine-tuning (SFT) workloads.

batch-level semanticsdata pipelinedistributed training

This study addresses the challenge of simultaneously achieving low latency, high throughput, and cost efficiency in enterprise-scale composite AI systems under concurrent heterogeneous model invocations. It presents the first systematic analysis of system-specific issues, including multi-model fan-out overhead, cascading cold-start propagation, and dynamic heterogeneity in scaling behavior. To tackle these challenges, the authors propose a modular, platform-agnostic inference architecture that integrates serverless computing, dynamic autoscaling, MLOps pipelines, and cooperative multi-model scheduling. Experimental results demonstrate that the proposed approach reduces P95 tail latency by over 50%, increases throughput by up to 3.9×, and achieves 30%–40% cost savings compared to baseline systems.

compound AI systemsheterogeneous model invocationslow-latency

This study addresses the significant gap between the static capabilities of open-source large language models and their real-world performance in hosted API services, a discrepancy exacerbated by insufficient understanding of service-layer heterogeneity and dynamics. Leveraging multidimensional data collected in Q4 2025 from the AI Ping platform—including request logs, metadata, compatibility probes, price snapshots, and latency measurements—the work uncovers three key patterns: demand concentration inertia, supply-demand misalignment, and task-conditioned routing. Building on these insights, the paper reframes model deployment as a constrained statistical decision problem and demonstrates that intelligent routing reduces inference costs for Qwen3-32B by 37.8% and increases throughput for DeepSeek-V3.2 by approximately 90%.

hosted APIsmodel deploymentopen-weight LLMs

Hot Scholars

PJ

Peng Jiang

Kuaishou Technology
Recommender SystemMachine LearningComputational Advertising
KG

Kun Gai

Senior Director & Researcher, Alibaba Group
Machine LearningComputational Advertising
RB

Ricardo Britto

Ericsson / Blekinge Institute of Technology
Software process improvementMachine LearningSearch-basedsoftware engineering
KK

Kirill Khrylchenko

Yandex
recommender systemsdeep learninguser modelingconversational recommenders