Score
Designs and implements systems and software that execute trained AI/ML models to produce predictions or decisions, including model runtimes, serving infrastructure, batching, quantization, and hardware-accelerated execution. Analyzes and optimizes inference workloads for latency, throughput, resource utilization, cost, reliability, and SLAs, and develops profiling, scheduling, and deployment strategies to meet application constraints.
The rapid advancement of generative AI (GenAI) has introduced unprecedented computational demands for both training and inference, necessitating systematic characterization of AI accelerator evolution. Method: This study comprehensively tracks publicly released AI accelerators from 2017 to 2024, proposing a novel taxonomy integrating dataflow, memory architecture, and workload characteristics. It constructs a standardized benchmark database covering peak compute, energy efficiency (TOPS/W), process node, and other key metrics—incorporating, for the first time, parameters of mainstream commercial GenAI chips. Using performance-power scatter plots, market-segmented scaling visualizations, and trend modeling, it analyzes technical trajectories and energy-efficiency bottlenecks across domains (e.g., cloud training, edge inference). Contribution/Results: The work establishes a sustainable, authoritative benchmark that provides quantitative foundations for AI hardware architecture design, industry roadmap planning, and evidence-based policy formulation.
This work addresses the challenge that existing AI systems struggle to dynamically observe, intervene in, and optimize agent behavior at runtime, making it difficult to simultaneously achieve high task success rates, low latency, token efficiency, reliability, and safety. To overcome this limitation, the paper proposes a novel runtime infrastructure layer situated between the model and the application, which treats the AI execution process itself as an optimizable object—departing from conventional approaches that restrict optimization to static model or log-level adjustments. This layer enables proactive intervention and multi-dimensional performance co-optimization through mechanisms such as runtime monitoring, real-time inference, adaptive memory management, fault recovery, and policy enforcement. Experimental results demonstrate that the proposed approach significantly enhances the holistic performance of long-horizon agent workflows across task success rate, response latency, token efficiency, system reliability, and safety compliance.
This work addresses the lack of systematic performance analysis in AI model deployment and inference, which hinders scalability and efficiency in real-world applications. Building upon BentoML, the study constructs a scalable inference system for a RoBERTa-based sentiment analysis model and identifies inference bottlenecks under three realistic traffic patterns: steady-state, bursty, and high-load scenarios. The authors propose a novel multi-level optimization framework tailored to practical deployment environments, applying coordinated improvements across runtime, service, and deployment layers. Leveraging statistical analysis, they quantify the impact of these optimizations and further evaluate the inference resilience of a single-node K3s cluster under perturbations. Experimental results demonstrate that the optimized system substantially reduces latency, increases throughput, and effectively enhances both the scalability and robustness of AI inference.
This work addresses the efficiency bottleneck in scientific discovery by machine learning agents caused by reliance on costly physical execution. To overcome this, the authors propose a prediction-first decision paradigm that internalizes execution priors, enabling large language models (LLMs) to perform rapid inference based on verified data analysis reports instead of time-consuming real-world trials. The study formalizes, for the first time, the data-driven solution preference task, constructs a large-scale pairwise comparison corpus, and introduces a prediction–verification loop mechanism alongside a world-model-inspired prediction architecture. This framework integrates Verified Data Analysis Report prompting with preference learning. Experiments demonstrate that the approach achieves a prediction accuracy of 61.5% with well-calibrated confidence, and the resulting FOREAGENT agent converges six times faster than baselines while outperforming execution-based methods by 6% in overall performance.
This paper addresses long-standing bottlenecks in traditional operating systems—namely, scalability, adaptability, and manageability—by proposing a systematic framework for deep AI–OS integration. Methodologically, it introduces the first three-stage AI–OS co-evolution path: AI-powered → AI-refactored → AI-driven; designs verifiable kernel-level inference mechanisms, hybrid rule- and AI-based decision models, and modular AI-ready kernel interfaces; and spans the full software stack (kernel, drivers, runtime, toolchain), integrating machine learning, large language models, and agent-based techniques—with real-time constraint modeling, dynamic workload forecasting, and edge-coordinated inference. Contributions include: (1) a unified analytical framework for AI–OS interaction; (2) standardized evaluation dimensions and a methodology pipeline; and (3) a production-oriented integration blueprint and benchmarking recommendations—collectively providing both theoretical foundations and actionable engineering pathways for AI–OS co-evolution.
Designing hardware platforms for distributed inference of next-generation large language models (LLMs) poses significant challenges due to architectural and optimization heterogeneity. This paper introduces GenZ, the first systematic software-hardware co-modeling framework for LLM inference. GenZ unifies analytical modeling of diverse architectures—including Dense, GQA, MoE, and Mamba—as well as key optimizations such as chunking, speculative decoding, and quantization. Leveraging analytical modeling, cycle-accurate simulation, and real-system calibration, it achieves end-to-end inference latency prediction with a geometric mean error of only 5.82%. GenZ precisely identifies dominant bottlenecks—computation, memory capacity/bandwidth, or network latency/bandwidth—under varying service-level objective (SLO) constraints, and enables multi-dimensional sensitivity analysis across architectures and optimization strategies. The framework is open-sourced and accompanied by an interactive web-based analysis platform.
Traditional AI systems rely on fixed monolithic models, which struggle to dynamically allocate resources, decompose tasks, or update knowledge in response to varying inputs, leading to degraded performance and increased costs. This work proposes the first system-level design methodology for distributed composite AI systems, formulating a design space through workflow topologies and configuration choices and identifying eight core design patterns. The framework jointly optimizes model selection and runtime parameters, enabling task decomposition, multi-model orchestration, and explicit control logic, thereby facilitating a shift from static monolithic architectures toward dynamic, composable, and adaptive ones. Evaluated across three case studies, the approach reduces latency by up to 60% and cost by up to 71%, with only a 2.5–4 percentage point drop in accuracy.
This work addresses the inefficiency of the current Python machine learning ecosystem in supporting large-scale, highly concurrent machine learning pipeline searches driven by large language model (LLM) agents. To overcome this limitation, we propose the first system architecture specifically designed for agent-driven ML workloads, which decouples the agent’s planning and reasoning from pipeline execution to enable batch compilation and efficient scheduling. Our system introduces pipeline graph compilation, batched execution optimization, and a high-performance Rust runtime, while seamlessly integrating with mainstream Python libraries and supporting heterogeneous backends including CPUs and GPUs. Experimental results demonstrate that our approach achieves up to a 16.6× speedup on large-scale agent-driven ML pipeline search tasks.
This study addresses the stringent constraints on size, power, and computational resources faced by AI inference on resource-limited platforms such as small satellites. By conducting empirical characterization of quantized AI inference on Cortex-M-class processors using representative embedded vision neural networks, the work establishes the first measurement-based performance baseline for on-board embedded systems. It introduces an explicit multi-core/multi-device cooperative scheduling mechanism and integrates analysis of ALU/SIMD utilization with memory traffic to evaluate system behavior. Moving beyond conventional paradigms that rely on opaque OS-level scheduling, this research provides comparable latency and data-movement benchmarks for typical spaceborne processors like LEON and NOEL-V, thereby demonstrating the critical role of architecture-aware design and cooperative scheduling as key dimensions in optimizing embedded AI inference for satellite applications.
This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.