🤖 AI Summary
This work addresses the challenges of computational idleness caused by heterogeneous splitting points and the difficulty of batching requests with varying depths when deploying large language models in the cloud under resource-constrained edge devices. To this end, we propose an edge-cloud collaborative inference runtime system. Its core innovations include dynamic model partitioning, KV cache localization, a depth-aware dyForward executor, and a Depth-Synchronized Batching (DSB) algorithm, which collectively enable efficient shared computation for heterogeneous requests at their deepest common suffix. Experimental results demonstrate that the proposed system improves throughput by 275% over FIFO scheduling and by 48% over exact-match batching. Furthermore, it significantly reduces average session latency while eliminating additional GPU waiting overhead.
📝 Abstract
Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models (LLMs) are increasingly used to support these applications. Yet, due to their high resource demands, LLMs are mostly deployed in the cloud. Layer-wise edge-cloud inference lets resource-constrained edge devices contribute computation to LLMs they cannot host in full. However, heterogeneous split points introduce two coupled inefficiencies. First, edge execution and communication create idle gaps between cloud invocations. Second, requests arriving at different model depths cannot be conventionally batched. We present DySCo, a collaborative runtime that keeps KV caches local and introduces dyForward, a model-aware layer-range executor that runs configurable contiguous layer ranges from resident model shards without reloading weights. For multi-edge serving settings, we introduce depth-synchronized batching (DSB), which advances heterogeneous requests to the deepest cut and batches their common suffix. Experiments across heterogeneous devices, two model families, and local and wide-area links show that idle gaps increase the latency of subsequent GPU forward calls even when waiting time is excluded, adding up to 25 ms of additional cloud-side suffix latency per decoding step in our measurements. At an average concurrency of eight, DSB improves throughput by 275% over FIFO, 48% over exact-match batching, and 79% over round-robin interleaving while reducing mean per-session latency. Together, these results show that requests with different edge-cloud splits can reuse resident cloud weights and share batched suffix computation. The artifact repository for this work is publicly available at: https://github.com/Large-scale-Sustainable-Computing-LSC/dysco-artifact