Score
Designs and implements infrastructure and runtime components that expose trained models as reliable, low‑latency, and scalable services for inference. This includes building APIs/endpoints, request routing and batching, caching, versioning and rollback, autoscaling and load balancing, deployment automation, and monitoring/observability to meet latency, throughput, and availability requirements.
This work addresses the challenge that existing AI systems struggle to dynamically observe, intervene in, and optimize agent behavior at runtime, making it difficult to simultaneously achieve high task success rates, low latency, token efficiency, reliability, and safety. To overcome this limitation, the paper proposes a novel runtime infrastructure layer situated between the model and the application, which treats the AI execution process itself as an optimizable object—departing from conventional approaches that restrict optimization to static model or log-level adjustments. This layer enables proactive intervention and multi-dimensional performance co-optimization through mechanisms such as runtime monitoring, real-time inference, adaptive memory management, fault recovery, and policy enforcement. Experimental results demonstrate that the proposed approach significantly enhances the holistic performance of long-horizon agent workflows across task success rate, response latency, token efficiency, system reliability, and safety compliance.
To address challenges in high-performance computing (HPC) environments—including heterogeneous LLM deployment, inflexible resource scheduling, and significant performance volatility under multi-model concurrent inference—this paper proposes a scalable LLM inference engine architecture built atop SLURM. The architecture integrates containerized microservices with dynamic resource orchestration, enabling fine-grained, coordinated allocation of CPU, GPU, and memory resources, and provides unified access via RESTful APIs to support both batch and interactive inference workloads. A novel multi-step “tribunal” refinement workflow is introduced to enhance fault tolerance and operational flexibility. Experiments on Llama-series models across multi-node HPC clusters demonstrate sub-50 ms latency and 128 concurrent requests for smaller models (e.g., Llama-3-8B), and stable dual-concurrent execution for large models (e.g., Llama-3-70B), with low scheduling overhead and strong horizontal scalability. The system has been successfully deployed in production applications, including retrieval-augmented generation chatbots.
In large language model (LLM) serving, autoscaling faces a fundamental trade-off between high scaling latency and substantial parameter caching overhead: conventional approaches rely on local parameter caches, causing service interruption during model loading and suffering from data-plane bottlenecks in cross-host scaling. This paper proposes a local-cache-free, millisecond-scale real-time autoscaling framework. It introduces the first O(1)-complexity network-based direct parameter transmission mechanism, enabling zero-copy parameter loading over GPU-to-GPU high-speed interconnects. We design a layer-granularity dynamic collaborative execution architecture supporting fine-grained load migration and multicast-based parameter distribution. Integrated cooperative inference scheduling eliminates cold-start delays. Experiments show up to 86% reduction in tail latency, achieving performance close to the ideal full-host, full-parameter caching configuration—while completely eliminating local parameter storage overhead.
Current large language model inference systems lack the capability to dynamically adjust model parallelism topologies at runtime, necessitating service restarts under varying workloads and causing multi-minute disruptions, loss of KV cache, and substantial recomputation overhead. This work proposes ReMP, the first framework enabling online elastic reconfiguration of combined tensor and pipeline parallelism. By decoupling topology from execution state, designing a two-dimensional KV cache migration mechanism, and orchestrating an end-to-end reconfiguration pipeline, ReMP reduces topology switching latency to 1–7 seconds across 7B–70B models—orders of magnitude faster than restarting. This significantly improves time-to-first-token (TTFT), time per output token (TPOT), and throughput under dynamic workloads.
This study addresses the lack of systematic empirical investigation into the deployment of large language model (LLM) inference frameworks in real-world software systems. Through large-scale analysis of open-source projects, combining static code analysis with metadata mining, this work systematically characterizes the adoption patterns and integration strategies of prominent frameworks—including vLLM, SGLang, TensorRT-LLM, LMDeploy, and FlashInfer—and examines their relationships with model type, scale, modality, and deployment environment. The findings reveal that vLLM exhibits the highest adoption rate, while multi-framework co-deployment remains limited yet holds complementary potential. Framework selection is strongly driven by model characteristics and application scenarios, effectively supporting diverse system designs such as reinforcement learning inference, multimodal generation, and microservice architectures.
This work addresses the lack of systematic performance analysis in AI model deployment and inference, which hinders scalability and efficiency in real-world applications. Building upon BentoML, the study constructs a scalable inference system for a RoBERTa-based sentiment analysis model and identifies inference bottlenecks under three realistic traffic patterns: steady-state, bursty, and high-load scenarios. The authors propose a novel multi-level optimization framework tailored to practical deployment environments, applying coordinated improvements across runtime, service, and deployment layers. Leveraging statistical analysis, they quantify the impact of these optimizations and further evaluate the inference resilience of a single-node K3s cluster under perturbations. Experimental results demonstrate that the optimized system substantially reduces latency, increases throughput, and effectively enhances both the scalability and robustness of AI inference.
Existing large model training data pipelines struggle to simultaneously ensure batch semantics, fault isolation, and consistency. This work proposes an agent-free, object-storage-native training data plane featuring three core innovations: a Transactional Global Batch (TGB) abstraction that guarantees training consistency, a storage-layer-embedded garbage collection mechanism aligning producer states with distributed checkpoints, and a communication-free Decentralized Adaptive Commit (DAC) algorithm. Leveraging lakehouse ACID semantics in object storage and distributed checkpointing, the system achieves higher throughput than colocated loaders and Kafka, lower read latency, and full fault isolation across 64-GPU multimodal pretraining and supervised fine-tuning (SFT) workloads.
This study addresses the challenge of simultaneously achieving low latency, high throughput, and cost efficiency in enterprise-scale composite AI systems under concurrent heterogeneous model invocations. It presents the first systematic analysis of system-specific issues, including multi-model fan-out overhead, cascading cold-start propagation, and dynamic heterogeneity in scaling behavior. To tackle these challenges, the authors propose a modular, platform-agnostic inference architecture that integrates serverless computing, dynamic autoscaling, MLOps pipelines, and cooperative multi-model scheduling. Experimental results demonstrate that the proposed approach reduces P95 tail latency by over 50%, increases throughput by up to 3.9×, and achieves 30%–40% cost savings compared to baseline systems.
This study addresses the significant gap between the static capabilities of open-source large language models and their real-world performance in hosted API services, a discrepancy exacerbated by insufficient understanding of service-layer heterogeneity and dynamics. Leveraging multidimensional data collected in Q4 2025 from the AI Ping platform—including request logs, metadata, compatibility probes, price snapshots, and latency measurements—the work uncovers three key patterns: demand concentration inertia, supply-demand misalignment, and task-conditioned routing. Building on these insights, the paper reframes model deployment as a constrained statistical decision problem and demonstrates that intelligent routing reduces inference costs for Qwen3-32B by 37.8% and increases throughput for DeepSeek-V3.2 by approximately 90%.