AIR-LLM: Broadcasting AI Weights over Radio for Memory-Free Edge LLM Inference via RF Computing

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of deploying large language models (LLMs) on memory-constrained edge devices, where high energy consumption during weight loading remains a critical bottleneck. To this end, we propose a storage-free LLM inference framework based on radio frequency (RF) broadcasting. Specifically, model weights are transmitted via RF signals, and matrix multiplications are executed directly in the RF domain using an RF-based GEMV architecture. By integrating MIMO spatial multiplexing with efficient precoding techniques, the framework enables user-agnostic concurrent computation. Experimental results demonstrate that the proposed approach reduces energy consumption by 157.7× and transmission latency by 104.1× in a 20-user scenario, while incurring only a marginal 4.0% increase in perplexity. This work presents an efficient new paradigm for edge AI deployment.
📝 Abstract
Next-generation large language models (LLMs) are expanding from the cloud to ubiquitous edge devices. However, edge devices typically either lack the memory to store increasingly large LLM weights or, even with enough memory, spend unaffordable energy on loading the weights. This raises our question: can an edge device run an LLM without storing or loading its weights, but receive them over the air and consume them on the fly? Inspired by wireless broadcasting, we present AIR-LLM, an LLM inference architecture for edge devices, which is composed of: (i) a central radio (e.g., 5G base stations) that broadcasts the LLM weights into the air, and (ii) the edge user that receives the weights and completes the general matrix-vector multiplication (GEMV) of LLM inference directly in the radio frequency (RF) domain using RF mixers. To further shorten the airtime, AIR-LLM exploits MIMO spatial multiplexing and proposes an energy-efficient precoder-postcoder pair on the edge to calibrate its own wireless channel. Since the central radio stays user-unaware, AIR-LLM is user-scalable so that one broadcast serves unlimited users within its coverage. We implement AIR-LLM on the NVIDIA Sionna ray-traced channels of two real-world urban scenes and the profiling of a real RF mixer. With a WikiText-2 perplexity degradation of 4.0% on LLaMA-3.1-8B, AIR-LLM saves the energy by 157.7x/40.4x against the FP16 and weight-only quantization baselines; with 20 users, its airtime is 104.1x/26.0x shorter, respectively.
Problem

Research questions and friction points this paper is trying to address.

Edge LLM Inference
Memory Constraint
Energy Consumption
Weight Loading
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

RF Computing
Edge LLM Inference
Wireless Broadcasting
MIMO Spatial Multiplexing
Memory-Free
🔎 Similar Papers
No similar papers found.