On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance

📅 2024-03-25
🏛️ arXiv.org
📈 Citations: 2
✨ Influential: 0
📄 PDF
🤖 AI Summary
Operator composition and performance trade-offs remain open challenges in multi-tier edge AI deployment. Method: This paper systematically evaluates the black-box deployment efficacy of three operator classes—model partitioning, quantization, and early-exit—across mobile-edge-cloud architectures, conducting cross-device empirical studies on ResNet, ResNeXt, DUC, and FCN vision models. Contribution/Results: We first reveal that a quantization + early-exit hybrid strategy achieves optimal latency–accuracy trade-off at the edge with only marginal accuracy degradation (<2%). We propose a scenario-aware deployment decision framework grounded in resource constraints, network conditions, and input scale: (i) mobile–edge collaborative partitioning under tight resource budgets; (ii) cloud-only execution for small-scale inputs; and (iii) edge-first execution for large-scale inputs. Experimental results demonstrate that quantization alone better preserves accuracy, whereas hybrid strategies significantly improve edge energy efficiency.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionComputer Vision: Learning & Optimization for CVSearch and Optimization: Learning to Search

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsResponsible Web: Human-perceived consequences of algorithmic deployment on the webGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
Deciding what combination of operators to use across the Edge AI tiers to achieve specific latency and model performance requirements is an open question for MLOps engineers. This study aims to empirically assess the accuracy vs inference time trade-off of different black-box Edge AI deployment strategies, i.e., combinations of deployment operators and deployment tiers. In this paper, we conduct inference experiments involving 3 deployment operators (i.e., Partitioning, Quantization, Early Exit), 3 deployment tiers (i.e., Mobile, Edge, Cloud) and their combinations on four widely used Computer-Vision models to investigate the optimal strategies from the point of view of MLOps developers. Our findings suggest that Edge deployment using the hybrid Quantization + Early Exit operator could be preferred over non-hybrid operators (Quantization/Early Exit on Edge, Partition on Mobile-Edge) when faster latency is a concern at medium accuracy loss. However, when minimizing accuracy loss is a concern, MLOps engineers should prefer using only a Quantization operator on edge at a latency reduction or increase, respectively over the Early Exit/Partition (on edge/mobile-edge) and Quantized Early Exit (on edge) operators. In scenarios constrained by Mobile CPU/RAM resources, a preference for Partitioning across mobile and edge tiers is observed over mobile deployment. For models with smaller input data samples (such as FCN), a network-constrained cloud deployment can also be a better alternative than Mobile/Edge deployment and Partitioning strategies. For models with large input data samples (ResNet, ResNext, DUC), an edge tier having higher network/computational capabilities than Cloud/Mobile can be a more viable option than Partitioning and Mobile/Cloud deployment strategies.
Problem

Research questions and friction points this paper is trying to address.

Assessing latency vs accuracy trade-offs in Edge AI deployment strategies
Evaluating hybrid operator combinations for optimal Edge AI performance
Determining best deployment tiers for resource-constrained Mobile/Edge scenarios
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid Quantization + Early Exit for Edge
Quantization operator prioritizes accuracy on Edge
Partitioning across Mobile-Edge for resource constraints
🔎 Similar Papers