🤖 AI Summary
This study addresses hardware utilization bottlenecks in multi-agent LLM inference under unified memory architectures (UMA), arising from bus contention, speculative decoding variance, and tool-call stalls. We propose a cross-layer inference system that transcends graph compiler constraints to enable UMA-aware parallelism, leveraging asymmetric memory layouts and zero-copy tensor parallelism to saturate heterogeneous compute resources. Furthermore, we introduce dynamic draft budget allocation based on real-time sequence predictability, alongside asynchronous suspension and yielding mechanisms to eliminate agent stalls. Evaluated on the Apple M4 SoC, UMA-aware execution achieves a 1.29× speedup, while the complete system delivers a 1.77× speedup in high-latency scenarios.
📝 Abstract
Emerging multi-agent LLMs demand privacy-preserving edge deployment, yet current inference systems struggle with these collaborative workflows. Specifically, the memory-bound decode phase causes severe bus contention on unified memory architectures (UMA), paralyzing naive CPU-GPU co-execution. Furthermore, speculative decoding in multi-agent workloads faces extreme variance in drafting difficulty, alternating between complex reasoning and predictable structured generation. Compounded by frequent tool-induced stalls, this highly fragmented execution severely underutilizes hardware and defeats traditional static batching.
We present EdgeAgent, a cross-layer inference system explicitly co-designed for edge UMA and multi-agent workloads. At the micro-architectural level, it bypasses rigid graph-compiler constraints to enable zero-copy UMA-aware tensor parallelism, utilizing asymmetric memory layouts to fully saturate both CPU and GPU compute units. At the scheduling level, it dynamically allocates draft budgets based on real-time sequence predictability to bound bandwidth waste. Concurrently, an asynchronous suspend-and-yield mechanism actively evicts stalled agents, ensuring continuous hardware saturation during unpredictable tool invocations.
Extensive evaluations on an Apple M4 SoC demonstrate that the UMA-aware execution alone contributes a 1.29x speedup over batched speculative decoding. Adding the agent-aware scheduling lifts the full EdgeAgent system to a 1.77x speedup under extreme tool-use latencies.