EdgeDAE: Acceleration of Diffusion Action Experts for Real-Time Physical AI with Tiny VLAs on Edge FPGA-GPU Systems

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory bottlenecks and computational underutilization in edge GPUs caused by repeated parameter loading during Diffusion Action Expert (DAE) inference. To overcome these limitations, we propose an FPGA-GPU heterogeneous computing architecture that strategically partitions workloads based on computational characteristics: the Vision Transformer is offloaded to the GPU, while DAE inference is executed on the FPGA leveraging BRAM/URAM on-chip memory, thereby completely eliminating memory access bottlenecks. The design further incorporates deep optimizations through quantization, fixed-point arithmetic, and hardware-based random number generation. Compared with edge GPUs, the proposed system reduces latency by 52.5% and improves throughput by 2.1×, achieving an energy efficiency approximately 17 times that of the RTX 4090.
📝 Abstract
Physical AI models such as Vision-Language-Action (VLA) architectures enable generalist robotic policies through large-scale transformer backbones and diffusion-based action decoders. While edge GPU platforms excel at parallelizing the compute-intensive vision-transformer workloads, they exhibit fundamental limitations for the Diffusion Action Expert (DAE) module: the iterative denoising process requires repeated parameter loading from DRAM across multiple steps, resulting in memory-bound performance where the GPU's massive computational throughput remains underutilized. This mismatch between DAE's I/O-intensive characteristics and GPU's compute-centric architecture motivates a heterogeneous acceleration approach. This paper presents \textbf{EdgeDAE}, a heterogeneous FPGA-GPU system that strategically partitions workloads based on computational characteristics. We offload the perception-heavy vision-transformer to GPU while accelerating DAE inference on FPGA through complete on-chip parameter storage in BRAM/URAM. This architecture eliminates the memory bottleneck by co-designing quantization strategies, fixed-point arithmetic, and hardware-efficient random number generation for the FPGA fabric. Compared to an edge GPU baseline, EdgeDAE reduces end-to-end inference latency by 52.5\% for Octo-Small and 38.7\% for Octo-Base, with up to $2.10\times$ higher throughput; compared to a consumer GPU (RTX~4090), it achieves ${\sim}17\times$ higher energy efficiency.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Action Expert
Edge Computing
Memory Bottleneck
Vision-Language-Action
Real-time Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Heterogeneous FPGA-GPU Acceleration
Diffusion Action Expert
Vision-Language-Action Models
On-chip Parameter Storage
Hardware-Software Co-design