optimize inference systems

Designs, implements, and analyzes inference runtimes and pipelines that execute trained models with low end-to-end latency, high throughput, and constrained resource use across targets such as edge devices, GPUs, and distributed clusters. Work includes creating inference-time algorithms and adaptive control policies, memory and scheduling strategies, parallel/streaming architectures and sharding for distributed execution, profiling and benchmarking tools, and other deployment optimizations to accelerate inference and lower operational cost.

optimizeinferencesystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.94
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$213K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the fragmented and application-level implementation of preprocessing, accelerator invocation, and postprocessing in neural inference on microcontrollers, which lacks system-wide coordination. To overcome this, the authors propose abstracting the inference pipeline as an operating system primitive and introduce SynapticOS—a runtime built atop Zephyr—that enables deterministic execution with zero heap usage and a constant memory footprint (peak: 2,784 bytes) through static memory pools, frame-level reset mechanisms, and phase-order validation. Integrated with priority-based job scheduling (real-time, normal, and best-effort) and PowerQuad DSP optimizations—including self-calibrating FFT and Q15 matrix multiplication—the system achieves 215.8 FPS (4.63 ms per frame) for face detection on the NXP FRDM-MCXN947 platform, yielding a 6.7× speedup over QEMU with software floating-point while incurring only a 20.7 KB Flash overhead and passing all 99 test cases.

inference pipelinesmicrocontrollerneural inference

FIRST: Federated Inference Resource Scheduling Toolkit for Scientific AI Model Access

Oct 15, 2025
AT
Aditya Tanikanti
🏛️ Argonne National Laboratory | The University of Chicago | University of Illinois Chicago

High-performance computing (HPC) environments lack secure, scalable, and distributed AI inference services tailored for scientific computing. Method: This paper introduces the first federated large language model (LLM) inference platform designed specifically for HPC. It features a cluster-agnostic, OpenAI-compatible API supporting multiple backends (e.g., vLLM); integrates Globus Auth/Compute for cross-domain identity management and function-level scheduling; and proposes a novel “hot-node” retention mechanism with automated scaling to jointly optimize low-latency interactive inference and high-throughput batch processing. Contribution/Results: Deployed on production HPC systems, the platform generates over one billion tokens daily without reliance on commercial cloud infrastructure. It delivers cloud-like AI inference capabilities—complete with enhanced security, broad accessibility, and improved resource utilization—for scientific LLMs on HPC for the first time.

Enabling private AI inference across distributed HPC clustersProviding scalable access to diverse AI models on-premisesSupporting parallel inference workloads via federated scheduling

Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing

Jun 03, 2025
JS
Jinhao Sheng
🏛️ Beijing Normal University | BNU-HKBU United International College

This paper addresses the deployment of multi-model inference pipelines on resource-constrained edge devices. We propose an end-to-end adaptive configuration framework that, for the first time, explicitly incorporates device resource constraints into joint optimization decisions. Our method integrates residual feature extraction, LSTM-based workload forecasting, and policy-gradient reinforcement learning to jointly optimize QoS guarantees (e.g., latency and throughput), operational cost, and real-time adaptability. Evaluated on a real Kubernetes-based edge cluster, the framework achieves a 27% reduction in average inference latency, a 31% increase in throughput, a 22% decrease in deployment cost, and over 42% faster configuration decision-making for complex pipelines—outperforming state-of-the-art baselines. The core contributions are: (i) a resource-aware joint optimization model that unifies hardware constraints with pipeline scheduling and scaling decisions; and (ii) a lightweight, learning-driven configuration mechanism enabling efficient, online adaptation under dynamic edge conditions.

Address device resource constraints in pipeline configurationOptimize QoS and costs for edge multi-model inference pipelinesReduce decision-making time for complex edge pipelines

This study addresses the limitations of dense large language models in long-chain reasoning, where fragmented KV caches and parallelization inefficiencies undermine traditional prefill extension strategies. Through systematic evaluation of dense and Mixture-of-Experts models ranging from 8B to 671B parameters on GPU clusters, the work uncovers critical performance bottlenecks: a sharp drop in data parallelism efficiency due to cache fragmentation, a nonlinear scaling inflection point in tensor parallelism around 32B parameters, and fundamental differences between sparse and dense architectures in interconnect bandwidth utilization and routing latency. Guided by extensive empirical analysis, the authors propose an architecture decision framework tailored to the “inference cliff” phenomenon, establishing design principles for next-generation LLM inference infrastructure that substantially improve resource utilization and throughput efficiency.

capacity-bound regimeinference scalingKV-cache fragmentation

Latest Papers

What's happening recently
View more

This work addresses the trade-off between efficiency and performance in multi-agent large language model inference, where existing approaches lack a unified framework for modeling diverse parallelization strategies. The study systematically distinguishes and unifies two inference-time parallelism mechanisms: replica parallelism—enabling multi-path exploration at the task level—and structural parallelism—supporting concurrent execution within a single reasoning path. To harmonize these strategies under a coherent execution semantics, the authors propose TIPEX, a novel coordination framework. Experimental results on the GAIA benchmark demonstrate that TIPEX significantly improves accuracy while reducing latency, with the most pronounced gains observed on medium-difficulty tasks. Notably, the findings also reveal that excessive parallelism can be detrimental, underscoring the necessity of tailoring parallelization strategies to the specific characteristics of each task.

execution coordinationinference-time parallelismmulti-agent LLM systems

This work reveals a novel attack surface in edge-cloud collaborative dual-path distributed inference systems: malicious “oscillating burst” traffic can induce resource contention in the slow path, causing benign requests to time out and be dropped, thereby triggering an “accuracy collapse” wherein the system degrades to low-precision fast-path outputs. Notably, this attack requires no access to the model or data and operates solely through network-level interference, significantly degrading perceptual performance. Evaluated on a multi-object tracking simulation platform under autonomous driving scenarios, approximately 4,000 burst requests increased the p99 latency for benign users from 92 ms to 2 seconds, reduced average HOTA by 7.0 points, and caused nearly 50% accuracy loss on rare classes such as stop signs.

accuracy collapsedistributed inference pipelineslatency deadline

This work addresses the challenge of achieving an optimal trade-off among latency, energy consumption, and accuracy in dynamic machine learning on edge devices, where both data distribution shifts and resource fluctuations are prevalent. To this end, the authors propose a two-tier adaptive architecture: a global scheduler deploys a lightweight cascade of expert and general-purpose models that adheres to system constraints, while a local controller continuously monitors data drift and hardware conditions to dynamically activate or deactivate expert models for improved inference efficiency. The key contributions include a formalized budget-constrained cascaded model formulation and a hierarchical control mechanism, both validated on embedded platforms. Experimental results demonstrate that, under distribution shifts, the approach reduces per-inference latency by up to 2.45× and energy consumption by up to 2.86× compared to static baselines, with less than 4% accuracy degradation.

data driftdynamic inferenceedge computing

This work addresses the challenge of efficient large-model inference in resource-constrained distributed edge environments by proposing a multi-cluster collaborative inference framework. In this framework, devices within each cluster employ lightweight local models to extract features, which are then uploaded to multi-GPU edge servers for fusion. The approach jointly optimizes model pruning ratios, task scheduling, bandwidth allocation, and transmit power to minimize inference distortion under constraints on latency, energy consumption, and server capacity. By innovatively integrating rate-distortion theory with partial information decomposition, the study reveals the fundamental trade-off between model pruning and collaborative inference, and establishes a unified optimization mechanism tailored for multi-cluster edge intelligence networks. Experimental results demonstrate that the proposed framework significantly outperforms existing methods, achieving simultaneous improvements in inference accuracy and resource efficiency.

collaborative inferenceedge intelligencelarge AI models

Hot Scholars

EC

Enhong Chen

University of Science and Technology of China
data miningrecommender systemmachine learning
BC

Beidi Chen

Carnegie Mellon University
Machine Learning
KG

Kun Gai

Senior Director & Researcher, Alibaba Group
Machine LearningComputational Advertising
HW

Hongzhi Wang

IBM Almaden Research Center
Medical Image Analysis
FW

Furu Wei

Distinguished Scientist, Microsoft Research
Natural Language ProcessingArtificial IntelligenceGeneral AIGenerative AI