🤖 AI Summary
This study addresses the oversight of future resource contention and service commitments in model-parallel inference scheduling for dynamic edge systems by proposing the PROMISE framework. PROMISE treats the committed completion time (CCT) as an endogenous decision variable, achieving commitment-aware parallel inference scheduling through predictive rolling-horizon optimization that integrates stochastic tasks, node mobility, and privacy-aware mechanisms. To resolve cross-slot coupling challenges, the framework employs joint SD-ES mapping, progress-aware surrogate evaluation, and analytical CCT recovery techniques. Experimental results demonstrate that PROMISE delivers robust scheduling performance across scenarios of varying scales, while its practical feasibility is further validated on a Raspberry Pi platform.
📝 Abstract
Model-parallel inference over dynamic edge systems requires scheduling decisions that account for not only instantaneous resources but also future resource contention and reliable service commitments. Existing edge-inference designs, however, predominantly optimize performance metrics based on current or short-term system states, without explicitly coupling current assignments with future commitment fulfillment. To address this issue, we propose PROMISE, a predictive rolling-horizon framework for commitment-aware model-parallel inference under spatio-temporal edge dynamics. PROMISE jointly models stochastic task generation, mobility-induced communication variations, privacy-aware model partitioning, and load-dependent edge computing capability. We introduce committed completion time (CCT) as an endogenous service decision and formulate joint SD--ES mapping and CCT determination to balance commitment fulfillment and service reward. To address cross-timeslot coupling, PROMISE estimates future task arrivals and computational workloads over an adaptive horizon and embeds them into certainty-equivalent ES-state rollout. A progress-aware surrogate then evaluates feasible current-stage mappings, while the corresponding CCTs are analytically recovered from predicted completion times. Only the first-stage decisions are committed, and the optimization is repeated with newly observed states in a receding-horizon manner. Numerical experiments demonstrate robust scheduling performance under diverse system scales and workload dynamics, while Raspberry-Pi-based experiments validate the practical feasibility and key system characteristics of model-parallel edge inference.