🤖 AI Summary
This study addresses the memory bottlenecks and computational underutilization in edge GPUs caused by repeated parameter loading during Diffusion Action Expert (DAE) inference. To overcome these limitations, we propose an FPGA-GPU heterogeneous computing architecture that strategically partitions workloads based on computational characteristics: the Vision Transformer is offloaded to the GPU, while DAE inference is executed on the FPGA leveraging BRAM/URAM on-chip memory, thereby completely eliminating memory access bottlenecks. The design further incorporates deep optimizations through quantization, fixed-point arithmetic, and hardware-based random number generation. Compared with edge GPUs, the proposed system reduces latency by 52.5% and improves throughput by 2.1×, achieving an energy efficiency approximately 17 times that of the RTX 4090.
📝 Abstract
Physical AI models such as Vision-Language-Action (VLA) architectures enable generalist robotic policies through large-scale transformer backbones and diffusion-based action decoders. While edge GPU platforms excel at parallelizing the compute-intensive vision-transformer workloads, they exhibit fundamental limitations for the Diffusion Action Expert (DAE) module: the iterative denoising process requires repeated parameter loading from DRAM across multiple steps, resulting in memory-bound performance where the GPU's massive computational throughput remains underutilized. This mismatch between DAE's I/O-intensive characteristics and GPU's compute-centric architecture motivates a heterogeneous acceleration approach. This paper presents \textbf{EdgeDAE}, a heterogeneous FPGA-GPU system that strategically partitions workloads based on computational characteristics. We offload the perception-heavy vision-transformer to GPU while accelerating DAE inference on FPGA through complete on-chip parameter storage in BRAM/URAM. This architecture eliminates the memory bottleneck by co-designing quantization strategies, fixed-point arithmetic, and hardware-efficient random number generation for the FPGA fabric. Compared to an edge GPU baseline, EdgeDAE reduces end-to-end inference latency by 52.5\% for Octo-Small and 38.7\% for Octo-Base, with up to $2.10\times$ higher throughput; compared to a consumer GPU (RTX~4090), it achieves ${\sim}17\times$ higher energy efficiency.