Weaver: A System for AI-RAN Compute Sharing with Foundation Model Training

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of utilizing idle GPUs at AI-RAN base stations for large model training without compromising real-time communications. To this end, it proposes Weaver, a system that pioneers the analysis of micro- and macro-level characteristics of idle AI-RAN compute resources. Weaver incorporates a computation-aware MAC scheduler to guarantee RAN prioritization and introduces a two-tier elastic training mechanism to smooth workloads and accommodate heterogeneous resources. Experimental evaluations based on an O-RAN architecture demonstrate that Weaver increases available idle compute capacity by 4.9× and achieves 83% utilization. Furthermore, multi-site distributed training throughput improves by 2.1× to 3.7× over baselines, realizing effective co-optimization of communication and computation.
📝 Abstract
The emergence of AI-RAN infrastructure, which equips cell sites with GPU-accelerated hardware, creates an opportunity to colocate non-RAN workloads with primary RAN processing. We explore using this spare capacity for decentralized training of foundation models (FMs), one of the most compute-intensive AI workloads. We present the first characterization of spare GPU capacity in AI-RAN systems at both micro-scale--across transmission slots within a cell site--and macro-scale--across sites. Our analysis finds that 40-85% of GPU capacity is unused; although this capacity is temporally bursty at individual sites, it is spatially complementary across sites. To safely and efficiently harness these resources, we present Weaver, a system that opportunistically trains FMs alongside latency-critical RAN workloads without degrading RAN performance. Weaver adopts a RAN-first design: a spare-compute controller integrated into the MAC scheduler uses compute-aware scheduling to smooth RAN GPU demand and exposes more usable spare GPU capacity. A two-level elastic training framework then adapts to dynamic, heterogeneous spare capacity within and across sites. Experiments on an O-RAN-aligned system prototype show that Weaver creates up to 4.9x more usable spare compute and utilizes up to 83% of the available spare capacity. On a multi-site testbed, Weaver improves training throughput by 2.1-3.7x over baseline approaches.
Problem

Research questions and friction points this paper is trying to address.

AI-RAN
foundation model training
spare GPU capacity
compute sharing
decentralized training
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI-RAN
Foundation Model Training
Compute-aware Scheduling
Elastic Training
Spare GPU Capacity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.