Adaptive Orchestration for Inference of Large Foundation Models at the Edge

📅 2025-03-19
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing distributed large language model (LLM) inference in resource-constrained, highly heterogeneous, and dynamically evolving edge computing environments—such as Multi-access Edge Computing (MEC)—faces multiple challenges: volatile workloads, fluctuating bandwidth, node-level congestion, and real-time evolution of privacy constraints. To address these, this paper proposes the first adaptive sharding architecture tailored for LLM inference at the edge. It enables runtime capacity-aware node selection, operational-condition-driven dynamic model partition redistribution, and layer-granular real-time repartitioning. Leveraging dynamic workload modeling and a joint QoS-and-privacy-aware orchestration mechanism, the framework achieves co-optimization of low latency, high throughput, and strong privacy guarantees. Evaluated on realistic MEC deployments, our approach improves inference throughput by 42.3% and resource utilization by 35.7%, while rigorously satisfying heterogeneous QoS and privacy SLA requirements.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionSearch and Optimization: Distributed SearchConstraint Satisfaction and Optimization: Distributed CSP/Optimization

Application Category

Economics, Online Markets and Human Computation: Architectures and workflows that use LLMs for crowd workUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
Large Foundation Models (LFMs), including multi-modal and generative AI models, promise to unlock new capabilities for next-generation Edge AI applications. However, performing inference with LFMs in resource-constrained and heterogeneous edge environments presents significant challenges for workload orchestration. We propose a novel adaptive orchestration method and system tailored specifically for managing distributed inference workloads across multi-access edge computing (MEC) infrastructures. Our approach enhances traditional workload orchestration by introducing dynamic methods including: (1) adaptive workload distribution that selects optimal, inter-connected edge nodes based on runtime capacity profiling; (2) dynamic redistribution of LFM partitions as operational conditions evolve, and; (3) real-time reconfiguration (e.g., re-splitting) of LFM layers to balance performance and privacy requirements. Our proposed framework introduces an architecture for adaptive split inference, enabling real-time, QoS-aware management of inference workloads. We present a reference architecture, detail operational mechanisms, and demonstrate its application through various use cases in real-world scenarios.
Problem

Research questions and friction points this paper is trying to address.

Orchestrating Large Foundation Model inference in resource-constrained edge environments
Adapting split inference to dynamic workloads and network conditions
Balancing latency, throughput, and privacy in edge AI applications
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive split inference orchestration framework
Dynamic partition migration for fluctuating conditions
Real-time reconfiguration balancing QoS metrics
🔎 Similar Papers
No similar papers found.
F
Fernando Koch
Florida Atlantic University, USA
A
Aladin Djuhera
Technical University Munich, Germany
A
A. Binotto
Carl Zeiss AG, Germany