real-time inference

Designs and builds systems that run trained models to produce predictions or decisions within strict latency and throughput constraints, handling streaming or interactive inputs and the associated batching, scheduling, and resource management. Analyzes end-to-end inference latency, jitter, scalability, and reliability and implements techniques such as model optimization, quantization, pruning, compilation, hardware acceleration, and adaptive routing to meet real-time requirements.

real-timeinference

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of achieving an optimal trade-off among latency, energy consumption, and accuracy in dynamic machine learning on edge devices, where both data distribution shifts and resource fluctuations are prevalent. To this end, the authors propose a two-tier adaptive architecture: a global scheduler deploys a lightweight cascade of expert and general-purpose models that adheres to system constraints, while a local controller continuously monitors data drift and hardware conditions to dynamically activate or deactivate expert models for improved inference efficiency. The key contributions include a formalized budget-constrained cascaded model formulation and a hierarchical control mechanism, both validated on embedded platforms. Experimental results demonstrate that, under distribution shifts, the approach reduces per-inference latency by up to 2.45× and energy consumption by up to 2.86× compared to static baselines, with less than 4% accuracy degradation.

data driftdynamic inferenceedge computing

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Nov 20, 2024
JF
Jared Fernandez
🏛️ Meta | Carnegie Mellon University

Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.

Assessing diminishing returns in scaling accelerators for large model trainingEvaluating parallelization strategies to minimize distributed communication overheadOptimizing hardware configuration for efficient large-scale model training

Optimizing edge AI models on HPC systems with the edge in the loop

May 26, 2025
MA
Marcel Aach
🏛️ Jülich Supercomputing Centre | Flanders Make

Deploying compact AI models on edge devices—such as those used in additive manufacturing—faces a fundamental trade-off among high accuracy, low latency, and minimal memory footprint. To address this, we propose an “edge-in-the-loop” hardware-aware neural architecture search (NAS) paradigm. Our approach integrates real-time latency measurements from physical edge hardware deployed in Belgium into a NAS workflow orchestrated on a high-performance computing (HPC) platform in Germany, establishing a cross-regional collaborative optimization loop. It further incorporates distributed heterogeneous training and fine-tuning on the RAISE-LPBF dataset. Compared to manually designed baselines, the discovered architectures achieve an 8.8× speedup in inference latency and a 1.35× improvement in quality-related accuracy metrics. This work significantly advances the feasibility and practicality of end-to-end co-optimization for edge AI models.

Coupling edge devices with HPC for real-time trainingLeveraging hardware-aware NAS for efficient architecture searchOptimizing edge AI models for speed and accuracy

This work addresses the inherent trade-off between latency and throughput in deploying dense large language models (e.g., Llama-3.1-70B/405B), particularly when model size exceeds device memory, where parallelization strategies critically impact performance. The study systematically evaluates tensor parallelism (TP), pipeline parallelism (PP), and their hybrid configurations for single-node inference, revealing that TP is more effective at reducing latency while PP better enhances throughput. Building on this insight, the authors propose dynamically adjusting the TP–PP ratio to finely tune the latency–throughput trade-off. Experimental results demonstrate that this approach enables system optimization tailored to diverse service objectives—such as meeting strict SLA requirements or maximizing throughput—thereby offering clear architectural guidance for deploying dense LLMs in practice.

dense LLMinference deploymentlatency-throughput tradeoff

Latest Papers

What's happening recently
View more

This work addresses the practical challenges of deploying machine learning models in real-world settings, where heterogeneous data protocols, non-standard formats, and infrastructure constraints often necessitate redundant construction of integration pipelines. To overcome these issues, we propose SMOCS—a containerized, streaming ML system built on Apache Kafka—that decouples infrastructure from application logic through layered abstraction and employs a three-threaded agent architecture to separate data ingestion, online training, and real-time inference. The framework enables configuration-driven, no-code deployment, offering platform independence, fault isolation, and horizontal scalability, thereby significantly lowering the barrier to entry for domain experts. SMOCS has been open-sourced on the Jefferson Lab GitHub repository and demonstrates both continuous online learning capability and strong engineering practicality.

infrastructure constraintsintegration pipelinesmachine learning deployment

This work addresses the lack of systematic performance analysis in AI model deployment and inference, which hinders scalability and efficiency in real-world applications. Building upon BentoML, the study constructs a scalable inference system for a RoBERTa-based sentiment analysis model and identifies inference bottlenecks under three realistic traffic patterns: steady-state, bursty, and high-load scenarios. The authors propose a novel multi-level optimization framework tailored to practical deployment environments, applying coordinated improvements across runtime, service, and deployment layers. Leveraging statistical analysis, they quantify the impact of these optimizations and further evaluate the inference resilience of a single-node K3s cluster under perturbations. Experimental results demonstrate that the optimized system substantially reduces latency, increases throughput, and effectively enhances both the scalability and robustness of AI inference.

AI inferencemodel servingperformance analysis

Traditional AI systems rely on fixed monolithic models, which struggle to dynamically allocate resources, decompose tasks, or update knowledge in response to varying inputs, leading to degraded performance and increased costs. This work proposes the first system-level design methodology for distributed composite AI systems, formulating a design space through workflow topologies and configuration choices and identifying eight core design patterns. The framework jointly optimizes model selection and runtime parameters, enabling task decomposition, multi-model orchestration, and explicit control logic, thereby facilitating a shift from static monolithic architectures toward dynamic, composable, and adaptive ones. Evaluated across three case studies, the approach reduces latency by up to 60% and cost by up to 71%, with only a 2.5–4 percentage point drop in accuracy.

Compound AI SystemsDistributed AIModel-Centric Design

Existing large-scale model training systems struggle to flexibly compose diverse parallelization strategies, often relying on manual expert tuning and lacking generality. This work proposes a programmable distributed training system that enables users to declaratively specify composite parallelism strategies—such as data, pipeline, and expert parallelism—through model annotations and scheduling directives. These specifications are compiled via a unified intermediate representation (IR) into device-level execution plans, fully decoupling strategy definition from runtime execution over a global compute-communication DAG. The system is the first to support automatic compilation of user-defined composite strategies, matching the performance of established approaches like ZeRO while significantly improving both performance and memory efficiency in complex scenarios such as DeepSeek-V3’s DualPipe.

distributed trainingflexibilitymodel parallelism

This work addresses the significant degradation in inference throughput caused by GPU memory constraints when concurrently deploying multiple large language models on shared heterogeneous hardware, where resource scheduling, model offloading, and preemption become critical bottlenecks. Through empirical methodologies—including cross-platform performance profiling, layer-wise offloading experiments, and fine-grained decomposition of preemption overhead—the study systematically uncovers, for the first time, the nonlinear relationship between offloading and throughput decline. It further identifies model state reloading as the primary source of preemption overhead. The findings reveal that smaller models are more sensitive to reduced GPU residency, and that such overhead is jointly influenced by model architecture and hardware characteristics. These insights motivate a scheduler design that integrates model-specific sensitivity with data migration costs, offering crucial guidance for building efficient multi-model serving systems.

CPU-GPU offloadingGPU memory constraintsheterogeneous hardware

Hot Scholars

IK

Itzik Klein

University of Haifa
RoboticsInertial SensingData-Driven NavigationAUV
RM

Ralf Mikut

Karlsruhe Institute of Technology (Germany)
data miningimage processingcomputational intelligencezebrafish
ZL

Zhibin Li

Professor in School of Transportation, Southeast University
Intelligent Transportation SystemTraffic ControlTraffic SafetyTraffic Flow
JX

Jiaping Xiao

Nanyang Technological University
Cyber-Physical SystemsIntelligent SystemsMultirobot LearningArtificial Intelligence
MF

Mir Feroskhan

School of Mechanical and Aerospace Engineering, Nanyang Technological University
Flight Dynamics and ControleVTOLUnmanned Aerial Systems