EdgeAgent: Orchestrating On-Device LLM inference for End-User Multi-Agent Systems on CPU-GPU Unified Memory Architectures

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses hardware utilization bottlenecks in multi-agent LLM inference under unified memory architectures (UMA), arising from bus contention, speculative decoding variance, and tool-call stalls. We propose a cross-layer inference system that transcends graph compiler constraints to enable UMA-aware parallelism, leveraging asymmetric memory layouts and zero-copy tensor parallelism to saturate heterogeneous compute resources. Furthermore, we introduce dynamic draft budget allocation based on real-time sequence predictability, alongside asynchronous suspension and yielding mechanisms to eliminate agent stalls. Evaluated on the Apple M4 SoC, UMA-aware execution achieves a 1.29× speedup, while the complete system delivers a 1.77× speedup in high-latency scenarios.
📝 Abstract
Emerging multi-agent LLMs demand privacy-preserving edge deployment, yet current inference systems struggle with these collaborative workflows. Specifically, the memory-bound decode phase causes severe bus contention on unified memory architectures (UMA), paralyzing naive CPU-GPU co-execution. Furthermore, speculative decoding in multi-agent workloads faces extreme variance in drafting difficulty, alternating between complex reasoning and predictable structured generation. Compounded by frequent tool-induced stalls, this highly fragmented execution severely underutilizes hardware and defeats traditional static batching. We present EdgeAgent, a cross-layer inference system explicitly co-designed for edge UMA and multi-agent workloads. At the micro-architectural level, it bypasses rigid graph-compiler constraints to enable zero-copy UMA-aware tensor parallelism, utilizing asymmetric memory layouts to fully saturate both CPU and GPU compute units. At the scheduling level, it dynamically allocates draft budgets based on real-time sequence predictability to bound bandwidth waste. Concurrently, an asynchronous suspend-and-yield mechanism actively evicts stalled agents, ensuring continuous hardware saturation during unpredictable tool invocations. Extensive evaluations on an Apple M4 SoC demonstrate that the UMA-aware execution alone contributes a 1.29x speedup over batched speculative decoding. Adding the agent-aware scheduling lifts the full EdgeAgent system to a 1.77x speedup under extreme tool-use latencies.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent LLMs
Unified Memory Architecture
Edge Inference
Speculative Decoding
Hardware Underutilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Memory Architecture
Multi-Agent Systems
Speculative Decoding
Tensor Parallelism
Edge Inference
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuhai Long
Sun Yat-sen University, School of Computer Science and Engineering, Guangzhou, China
Y
Yuanxin Wei
Sun Yat-sen University, School of Computer Science and Engineering, Guangzhou, China
K
Kai Wu
China Mobile Internet Company Ltd., Guangzhou, China
J
Jinhui Wei
Sun Yat-sen University, School of Computer Science and Engineering, Guangzhou, China
Dan Huang
Dan Huang
Sun Yat-sen University
HPCAI SystemIO Subsystem
J
Jiangsu Du
Sun Yat-sen University, School of Computer Science and Engineering, Guangzhou, China