🤖 AI Summary
Existing distributed large language model (LLM) inference in resource-constrained, highly heterogeneous, and dynamically evolving edge computing environments—such as Multi-access Edge Computing (MEC)—faces multiple challenges: volatile workloads, fluctuating bandwidth, node-level congestion, and real-time evolution of privacy constraints. To address these, this paper proposes the first adaptive sharding architecture tailored for LLM inference at the edge. It enables runtime capacity-aware node selection, operational-condition-driven dynamic model partition redistribution, and layer-granular real-time repartitioning. Leveraging dynamic workload modeling and a joint QoS-and-privacy-aware orchestration mechanism, the framework achieves co-optimization of low latency, high throughput, and strong privacy guarantees. Evaluated on realistic MEC deployments, our approach improves inference throughput by 42.3% and resource utilization by 35.7%, while rigorously satisfying heterogeneous QoS and privacy SLA requirements.
📝 Abstract
Large Foundation Models (LFMs), including multi-modal and generative AI models, promise to unlock new capabilities for next-generation Edge AI applications. However, performing inference with LFMs in resource-constrained and heterogeneous edge environments presents significant challenges for workload orchestration. We propose a novel adaptive orchestration method and system tailored specifically for managing distributed inference workloads across multi-access edge computing (MEC) infrastructures. Our approach enhances traditional workload orchestration by introducing dynamic methods including: (1) adaptive workload distribution that selects optimal, inter-connected edge nodes based on runtime capacity profiling; (2) dynamic redistribution of LFM partitions as operational conditions evolve, and; (3) real-time reconfiguration (e.g., re-splitting) of LFM layers to balance performance and privacy requirements. Our proposed framework introduces an architecture for adaptive split inference, enabling real-time, QoS-aware management of inference workloads. We present a reference architecture, detail operational mechanisms, and demonstrate its application through various use cases in real-world scenarios.