Going Beyond the Edge: Distributed Inference of Transformer Models on Ultra-Low-Power Wireless Devices

📅 Unknown Date
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of deploying large Transformer models on ultra-low-power Internet-of-Things (IoT) devices, which is hindered by their substantial computational and memory demands. To overcome this limitation, the authors propose CATS, a framework that enables distributed inference across multiple resource-constrained devices through co-design of model partitioning, communication mechanisms, and training strategies. Key innovations include SomeGather—a novel sparse communication primitive—communication-aware model parallelism, and a message-dropout technique for robust training under unreliable communication. These advances collectively reduce both communication overhead and memory footprint while preserving model accuracy. The framework demonstrates, for the first time, successful execution of a Transformer model 14× larger than what a single device can support, scaling up to 16 cooperating wireless edge devices.
📝 Abstract
Transformer models are rapidly becoming a cornerstone of modern Internet of Things (IoT) applications, yet their computational and memory demands far exceed the capabilities of a single typical ultra-low-power IoT device. We present CATS, a framework for distributed transformer inference on ultra-low-power wireless devices, enabling multiple devices to collaboratively execute models far larger than what a single device can sustain. At its core, CATS is a communication-aware distributed transformer inference scheme co-designed across transformer partitioning, wireless communication and training. It employs SomeGather, a new pruned communication primitive that selectively broadcasts activation columns to reduce communication bandwidth and RAM usage without sacrificing model accuracy. Building on SomeGather, we design a partitioning method that exploits this primitive for efficient model parallelism. To cope with unreliable wireless communication, CATS employs message-dropout during training, which mimics packet losses and yields models that are robust to message loss during inference. In real-world experiments, we show that CATS brings distributed transformer inference to ultra-low-power wireless devices for the first time, with deployments on up to 16 devices that collaboratively execute transformer models up to 14 times larger than what a single device can run.
Problem

Research questions and friction points this paper is trying to address.

Transformer models
ultra-low-power devices
distributed inference
IoT
model parallelism
Innovation

Methods, ideas, or system contributions that make the work stand out.

distributed inference
ultra-low-power IoT
SomeGather
communication-aware partitioning
message-dropout
💼 Related Jobs
No related jobs found.
A
Alexander Gräfe
RWTH Aachen University
D
Ding Huo
TU Darmstadt
V
Vincent de Bakker
RWTH Aachen University
J
Johannes Berger
RWTH Aachen University
Marco Zimmerling
Marco Zimmerling
Professor, TU Darmstadt
Embedded SystemsCyber-Physical SystemsWireless NetworkingInternet of Things
Sebastian Trimpe
Sebastian Trimpe
Professor, RWTH Aachen University
ControlMachine LearningNetworked SystemsRobotics