🤖 AI Summary
This work addresses the challenge of deploying large Transformer models on ultra-low-power Internet-of-Things (IoT) devices, which is hindered by their substantial computational and memory demands. To overcome this limitation, the authors propose CATS, a framework that enables distributed inference across multiple resource-constrained devices through co-design of model partitioning, communication mechanisms, and training strategies. Key innovations include SomeGather—a novel sparse communication primitive—communication-aware model parallelism, and a message-dropout technique for robust training under unreliable communication. These advances collectively reduce both communication overhead and memory footprint while preserving model accuracy. The framework demonstrates, for the first time, successful execution of a Transformer model 14× larger than what a single device can support, scaling up to 16 cooperating wireless edge devices.
📝 Abstract
Transformer models are rapidly becoming a cornerstone of modern Internet of Things (IoT) applications, yet their computational and memory demands far exceed the capabilities of a single typical ultra-low-power IoT device. We present CATS, a framework for distributed transformer inference on ultra-low-power wireless devices, enabling multiple devices to collaboratively execute models far larger than what a single device can sustain. At its core, CATS is a communication-aware distributed transformer inference scheme co-designed across transformer partitioning, wireless communication and training. It employs SomeGather, a new pruned communication primitive that selectively broadcasts activation columns to reduce communication bandwidth and RAM usage without sacrificing model accuracy. Building on SomeGather, we design a partitioning method that exploits this primitive for efficient model parallelism. To cope with unreliable wireless communication, CATS employs message-dropout during training, which mimics packet losses and yields models that are robust to message loss during inference. In real-world experiments, we show that CATS brings distributed transformer inference to ultra-low-power wireless devices for the first time, with deployments on up to 16 devices that collaboratively execute transformer models up to 14 times larger than what a single device can run.