Risk-Aware and Stable Edge Server Selection Under Network Latency SLOs

📅 2026-04-23
📈 Citations: 0
Influential: 0
📄 PDF

career value

225K/year
🤖 AI Summary
This work addresses the challenge of dynamic server selection in edge computing under stringent latency SLOs, where decisions must jointly manage tail-risk violations and switching stability. The authors propose a lightweight, interpretable decision framework that uniquely co-optimizes tail-risk control and switching stability: it estimates SLO violation risk using normal approximation and the Cantelli inequality, while incorporating a hysteresis mechanism to suppress excessive switching. Experimental results under a 0.5-second SLO demonstrate that, compared to a baseline relying solely on mean latency, the proposed method reduces deadline miss rate from 39% to 34%, cuts switching frequency by 88% (down to 5.5%), and maintains average latency stably at 0.45 seconds, thereby significantly enhancing both system robustness and efficiency.

Technology Category

Application Category

📝 Abstract
We present a lightweight and interpretable decision framework for dynamic edge server selection in latency-critical applications that explicitly accounts for tail risk and switching stability. Each candidate server is characterised by predictive mean and uncertainty summaries of network latency, which are used to estimate the risk of service-level objective (SLO) violations and to guide selection. Risk is evaluated using a tight Normal approximation complemented by a conservative Cantelli bound, while percentile-based scoring coupled with hysteresis stabilizes decisions and suppresses oscillatory switching under short-lived network fluctuations. Experimental results on a multi-server edge testbed with a strict SLO of $τ= 0.5$\,s show that the proposed approach reduces the deadline-miss rate from 39\% to 34\% compared to a mean-only baseline, while reducing switching frequency from 46\% to 5.5\% ($\approx$88\% reduction) and maintaining sub-SLO average latency ($\approx$0.45\,s). These results demonstrate that explicit risk evaluation combined with stability-preserving control enables practical and robust adaptive server selection in dynamic edge environments.
Problem

Research questions and friction points this paper is trying to address.

edge server selection
network latency
service-level objective (SLO)
risk-awareness
switching stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

risk-aware server selection
latency SLO
tail risk
hysteresis control
edge computing