🤖 AI Summary
This work addresses the significant memory bottlenecks encountered when deploying large Transformer models on memory-constrained Android devices (3.3–7.4 GB RAM). The authors propose CROWDio, a system enabling efficient distributed inference without requiring modifications to ONNX models. Its key innovations include just-in-time deferred partition loading, single-partition residency constraints, a four-level affinity scheduler, zlib-compressed tensor transmission, and stream-based 1:1 dependency scheduling. CROWDio achieves the first successful execution of DistilBERT on resource-limited Android clusters, with per-device peak memory consumption of only 43 ± 2 MB and energy usage of 50 ± 3 mAh per inference. Compared to barrier synchronization, its streaming concurrent batching reduces latency by 34%.
📝 Abstract
Deploying large deep neural networks on memory-constrained mobile devices is a central challenge in edge ML. While compression, pruning, and quantization reduce per-parameter cost, transformer-based models remain too large for the 3.3-7.4 GB RAM envelope of commodity Android handsets. We present the DNN pipeline scheduling subsystem of CROWDio, which achieves practical ONNX inference across resource-constrained Android workers without model modification, by distributing memory pressure across devices via five mechanisms: JIT deferred partition loading, a single-partition-resident constraint, a 4-tier affinity scheduler, a zlib-compressed tensor transport, and a streaming 1:1 dependency model. Evaluated on DistilBERT (Sanh et al., 2019) (approximately 67 M parameters, SST-2) across five Android handsets over ten runs, our system holds peak per-device RSS to 43+-2 MB and limits battery draw to 50+-3 mAh per run, while streaming concurrency cuts batch latency 34% below barrier synchronisation.