model inference

Designs and implements systems and software that execute trained AI/ML models to produce predictions or decisions, including model runtimes, serving infrastructure, batching, quantization, and hardware-accelerated execution. Analyzes and optimizes inference workloads for latency, throughput, resource utilization, cost, reliability, and SLAs, and develops profiling, scheduling, and deployment strategies to meet application constraints.

modelinference

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge that existing AI systems struggle to dynamically observe, intervene in, and optimize agent behavior at runtime, making it difficult to simultaneously achieve high task success rates, low latency, token efficiency, reliability, and safety. To overcome this limitation, the paper proposes a novel runtime infrastructure layer situated between the model and the application, which treats the AI execution process itself as an optimizable object—departing from conventional approaches that restrict optimization to static model or log-level adjustments. This layer enables proactive intervention and multi-dimensional performance co-optimization through mechanisms such as runtime monitoring, real-time inference, adaptive memory management, fault recovery, and policy enforcement. Experimental results demonstrate that the proposed approach significantly enhances the holistic performance of long-horizon agent workflows across task success rate, response latency, token efficiency, system reliability, and safety compliance.

Agent ExecutionAI RuntimeLong-horizon Workflows

This work addresses the lack of systematic performance analysis in AI model deployment and inference, which hinders scalability and efficiency in real-world applications. Building upon BentoML, the study constructs a scalable inference system for a RoBERTa-based sentiment analysis model and identifies inference bottlenecks under three realistic traffic patterns: steady-state, bursty, and high-load scenarios. The authors propose a novel multi-level optimization framework tailored to practical deployment environments, applying coordinated improvements across runtime, service, and deployment layers. Leveraging statistical analysis, they quantify the impact of these optimizations and further evaluate the inference resilience of a single-node K3s cluster under perturbations. Experimental results demonstrate that the optimized system substantially reduces latency, increases throughput, and effectively enhances both the scalability and robustness of AI inference.

AI inferencemodel servingperformance analysis

This work addresses the efficiency bottleneck in scientific discovery by machine learning agents caused by reliance on costly physical execution. To overcome this, the authors propose a prediction-first decision paradigm that internalizes execution priors, enabling large language models (LLMs) to perform rapid inference based on verified data analysis reports instead of time-consuming real-world trials. The study formalizes, for the first time, the data-driven solution preference task, constructs a large-scale pairwise comparison corpus, and introduces a prediction–verification loop mechanism alongside a world-model-inspired prediction architecture. This framework integrates Verified Data Analysis Report prompting with preference learning. Experiments demonstrate that the approach achieves a prediction accuracy of 61.5% with well-calibrated confidence, and the resulting FOREAGENT agent converges six times faster than baselines while outperforming execution-based methods by 6% in overall performance.

Data-centric Solution PreferenceExecution BottleneckHypothesis Evaluation

Integrating Artificial Intelligence into Operating Systems: A Survey on Techniques, Applications, and Future Directions

Jul 19, 2024
YZ
Yifan Zhang
🏛️ Zhejiang University | State Key Laboratory of Mathematical Engineering and Advanced Computing | Chinese Academy of Engineering

This paper addresses long-standing bottlenecks in traditional operating systems—namely, scalability, adaptability, and manageability—by proposing a systematic framework for deep AI–OS integration. Methodologically, it introduces the first three-stage AI–OS co-evolution path: AI-powered → AI-refactored → AI-driven; designs verifiable kernel-level inference mechanisms, hybrid rule- and AI-based decision models, and modular AI-ready kernel interfaces; and spans the full software stack (kernel, drivers, runtime, toolchain), integrating machine learning, large language models, and agent-based techniques—with real-time constraint modeling, dynamic workload forecasting, and edge-coordinated inference. Contributions include: (1) a unified analytical framework for AI–OS interaction; (2) standardized evaluation dimensions and a methodology pipeline; and (3) a production-oriented integration blueprint and benchmarking recommendations—collectively providing both theoretical foundations and actionable engineering pathways for AI–OS co-evolution.

Addressing OS bottlenecks in scalability, adaptability, and manageabilityDeveloping AI-enhanced systems to replace heuristic-based OS designsIntegrating AI techniques like ML and LLMs into operating systems

Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Jun 03, 2024
AB
A. Bambhaniya
🏛️ Georgia Institute of Technology | Intel Labs

Designing hardware platforms for distributed inference of next-generation large language models (LLMs) poses significant challenges due to architectural and optimization heterogeneity. This paper introduces GenZ, the first systematic software-hardware co-modeling framework for LLM inference. GenZ unifies analytical modeling of diverse architectures—including Dense, GQA, MoE, and Mamba—as well as key optimizations such as chunking, speculative decoding, and quantization. Leveraging analytical modeling, cycle-accurate simulation, and real-system calibration, it achieves end-to-end inference latency prediction with a geometric mean error of only 5.82%. GenZ precisely identifies dominant bottlenecks—computation, memory capacity/bandwidth, or network latency/bandwidth—under varying service-level objective (SLO) constraints, and enables multi-dimensional sensitivity analysis across architectures and optimization strategies. The framework is open-sourced and accompanied by an interactive web-based analysis platform.

Balancing LLM architectures and serving optimizations with platform parametersDetermining compute, memory, and network requirements for diverse LLM use casesEfficient hardware platform design for next-gen LLM inference

Latest Papers

What's happening recently
View more

Traditional AI systems rely on fixed monolithic models, which struggle to dynamically allocate resources, decompose tasks, or update knowledge in response to varying inputs, leading to degraded performance and increased costs. This work proposes the first system-level design methodology for distributed composite AI systems, formulating a design space through workflow topologies and configuration choices and identifying eight core design patterns. The framework jointly optimizes model selection and runtime parameters, enabling task decomposition, multi-model orchestration, and explicit control logic, thereby facilitating a shift from static monolithic architectures toward dynamic, composable, and adaptive ones. Evaluated across three case studies, the approach reduces latency by up to 60% and cost by up to 71%, with only a 2.5–4 percentage point drop in accuracy.

Compound AI SystemsDistributed AIModel-Centric Design

This work addresses the inefficiency of the current Python machine learning ecosystem in supporting large-scale, highly concurrent machine learning pipeline searches driven by large language model (LLM) agents. To overcome this limitation, we propose the first system architecture specifically designed for agent-driven ML workloads, which decouples the agent’s planning and reasoning from pipeline execution to enable batch compilation and efficient scheduling. Our system introduces pipeline graph compilation, batched execution optimization, and a high-performance Rust runtime, while seamlessly integrating with mainstream Python libraries and supporting heterogeneous backends including CPUs and GPUs. Experimental results demonstrate that our approach achieves up to a 16.6× speedup on large-scale agent-driven ML pipeline search tasks.

agentic pipeline searchlarge language modelsmassive ML workloads

This study addresses the stringent constraints on size, power, and computational resources faced by AI inference on resource-limited platforms such as small satellites. By conducting empirical characterization of quantized AI inference on Cortex-M-class processors using representative embedded vision neural networks, the work establishes the first measurement-based performance baseline for on-board embedded systems. It introduces an explicit multi-core/multi-device cooperative scheduling mechanism and integrates analysis of ALU/SIMD utilization with memory traffic to evaluate system behavior. Moving beyond conventional paradigms that rely on opaque OS-level scheduling, this research provides comparable latency and data-movement benchmarks for typical spaceborne processors like LEON and NOEL-V, thereby demonstrating the critical role of architecture-aware design and cooperative scheduling as key dimensions in optimizing embedded AI inference for satellite applications.

embedded platformsonboard computingquantized AI inference

This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.

capacity planningload testingML model serving

Hot Scholars

XJ

Xue Jiang

Peking University
Program GenerationLLM4SE
YD

Yihong Dong

Peking University
Code GenerationLarge Language Models
YK

Youngbin Kim

Senior Researcher, ETRI (Electronics and Telecommunications Research Institute)
RH

Ruida Hu

Harbin Institute of Technology, Shenzhen
software engineeringLLM agent
AC

Aman Chadha

GenAI Leadership @ Apple • Stanford AI • UW-Madison ECE • Ex: Apple, AWS, Alexa, Nvidia
Multimodal AINatural Language ProcessingComputer VisionSpeech Processing