real-time inference orchestration

Designs and implements systems, pipelines, and deployment artifacts that orchestrate low-latency real-time model inference across services and devices (including embedded targets), handling both synchronous and asynchronous execution modes. This work includes building and analyzing execution and deployment strategies—such as real-time chunking, temporal ensembling, and unified strategy/configuration management—to meet latency, resource, and consistency requirements.

real-timeinferenceorchestration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$193K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenges of high latency, unstable concurrency, and security risks faced by large language model (LLM) agents in automating asset lifecycle management within Industry 4.0. The authors propose a Plan-then-Execute architecture that generates verifiable workflow graphs and integrates a topology-aware parallel scheduling mechanism to enable controlled inference overlap while ensuring functional correctness and security. Key technical contributions include topological-sort-based multi-agent scheduling, structured context pruning, dependency-aware concurrency control, and graceful degradation under fault injection. Evaluated on the AssetOpsBench benchmark, the system reduces median end-to-end latency by 1.6× (up to 1.8× for highly parallel tasks) and cuts inference overhead by approximately 30% through context pruning, all while maintaining stable task completion rates and output quality.

concurrency instabilityIndustry 4.0latency

Joint Partitioning and Placement of Foundation Models for Real-Time Edge AI

Nov 30, 2025
AD
Aladin Djuhera
🏛️ Technical University of Munich | Florida Atlantic University | Carl Zeiss AG

To address resource dynamics, multi-objective constraints (latency, utilization, privacy), and infrastructure instability in large language model (LLM) inference within heterogeneous edge environments, this paper proposes a runtime-reconfigurable framework for joint model partitioning and deployment optimization. It is the first to formulate layer-granular model partitioning and device placement as a dynamic constrained optimization problem, integrating model-aware capacity analysis, dynamic graph neural network–based repartitioning, and resource forecasting. Evaluated in a 6G multi-access edge computing scenario, the approach reduces end-to-end latency by 27.4% and improves average GPU utilization by 39.1% over static baselines, while enabling privacy-sensitive layers to execute locally. The core contribution lies in an online, constraint-adaptive inference scheduler that ensures theoretical rigor and practical deployability under time-varying operational conditions.

Addressing resource volatility in heterogeneous edge environmentsDynamically partitioning and placing foundation models for edge AI inferenceOptimizing real-time inference under latency, utilization, and privacy constraints

Current large language model inference systems lack the capability to dynamically adjust model parallelism topologies at runtime, necessitating service restarts under varying workloads and causing multi-minute disruptions, loss of KV cache, and substantial recomputation overhead. This work proposes ReMP, the first framework enabling online elastic reconfiguration of combined tensor and pipeline parallelism. By decoupling topology from execution state, designing a two-dimensional KV cache migration mechanism, and orchestrating an end-to-end reconfiguration pipeline, ReMP reduces topology switching latency to 1–7 seconds across 7B–70B models—orders of magnitude faster than restarting. This significantly improves time-to-first-token (TTFT), time per output token (TPOT), and throughput under dynamic workloads.

DowntimeKV CacheLLM Serving

This work addresses the high latency in multi-agent systems caused by multi-step reasoning and redundant agent invocations during parallel execution, which often fails to meet real-time requirements. To this end, the authors propose LAMaS, a novel framework that introduces explicit latency supervision signals into multi-agent orchestration for the first time. LAMaS employs a learning-driven controller to construct an execution topology graph and leverages critical path analysis to optimize parallel scheduling. This approach departs from conventional paradigms centered on task performance or cost, instead prioritizing latency reduction along the critical path. Experimental results demonstrate that LAMaS reduces critical path length by 38%–46% compared to state-of-the-art methods across multiple benchmarks, while maintaining or even improving task performance.

inference latencylatencymulti-agent systems

Latest Papers

What's happening recently
View more

Traditional AI systems rely on fixed monolithic models, which struggle to dynamically allocate resources, decompose tasks, or update knowledge in response to varying inputs, leading to degraded performance and increased costs. This work proposes the first system-level design methodology for distributed composite AI systems, formulating a design space through workflow topologies and configuration choices and identifying eight core design patterns. The framework jointly optimizes model selection and runtime parameters, enabling task decomposition, multi-model orchestration, and explicit control logic, thereby facilitating a shift from static monolithic architectures toward dynamic, composable, and adaptive ones. Evaluated across three case studies, the approach reduces latency by up to 60% and cost by up to 71%, with only a 2.5–4 percentage point drop in accuracy.

Compound AI SystemsDistributed AIModel-Centric Design

This work addresses the high latency faced by industrial agents on the AssetOpsBench benchmark, which stems from repeated overhead in tool discovery, LLM-based planning, and MCP execution. Traditional semantic caching proves ineffective in scenarios sensitive to time, asset, or sensor parameters. To overcome this limitation, the authors propose a time-aware semantic caching mechanism that integrates disk-backed tool discovery caching with a dependency-aware parallel execution strategy, thereby optimizing the plan–execute pipeline. Experimental results demonstrate that the approach reduces end-to-end median latency by approximately 40% and accelerates MCP workflows by up to 1.67×. Notably, when time-aware cache hits occur, median speedups reach as high as 30.6×, underscoring the inherent limitations of purely semantic caching in complex industrial queries.

Industrial AgentsLLM CachingPlan-Execute Pipeline

This work addresses the orchestration bottlenecks faced by ultra-large-scale Sim-AI workflows on leadership-class supercomputers, which arise from task heterogeneity and extreme ensemble sizes. To overcome these challenges, the authors propose EnsembleLauncher, a recursively hierarchical and fully decentralized workflow orchestrator that introduces a decentralized control plane and a programmable scheduling policy interface, thereby surpassing conventional tools in both scalability and scheduling flexibility. Experiments on the Aurora supercomputer demonstrate that EnsembleLauncher can efficiently schedule system-wide resources to support up to 8 million serial tasks, achieving more than a fourfold performance improvement over state-of-the-art alternatives. Furthermore, it significantly enhances resource utilization for workloads with high task variance and active learning pipelines.

exascaleorchestration bottlenecksscalability

Hot Scholars

KH

Kaibin Huang

Professor and Dept.Head, University of Hong Kong; NAI Fellow; IEEE Fellow; Highly Cited Researcher
Machine LearningMobile Edge ComputingWireless CommunicationsWireless Power Transfer
JL

Jiaoyang Li

Assistant Professor at Robotics Institute, Carnegie Mellon University
Artificial IntelligenceMulti-Agent/Robot SystemsHeuristic SearchAutomated Planning
XC

Xianhao Chen

Assistant Professor, The University of Hong Kong
Wireless networksmobile edge computingedge AIdistributed learning
ZW

Zhanwei Wang

The University of Hong Kong
Edge IntelligenceWireless Communication
JM

Juan M. Lavista Ferres

Chief Scientist and Lab Director, Microsoft AI for Good Research Lab
Medical ImagingDeep LearningCausalityMachine Learning