Beyond LLM Serving: Characterizing Vision-Language-Action Workloads for Embodied AI System Design

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of real-time control and batched inference bottlenecks encountered when deploying Vision-Language-Action (VLA) models for embodied intelligence on edge devices. We systematically characterize the runtime behavior of four VLA models on edge GPUs and SoCs through single-inference profiling, closed-loop simulation, and hardware frequency scaling strategies. For the first time, this work reveals the transition mechanisms between memory and compute bottlenecks, demonstrating that no Pareto-optimal configuration exists, and proposes a theoretical framework for accuracy-speed-energy trade-offs. Furthermore, we quantify the trade-off effects of overlapping inference and execution. These findings provide critical guidance for the co-design of VLA model architectures, hardware selection, and runtime optimization strategies in resource-constrained edge environments.
📝 Abstract
Vision-language-action (VLA) models translate multimodal observations into low-level robot actions. During robot operation, each control period sets an inference deadline, and overruns leave the robot acting on stale observations, reducing task success. Meeting this deadline motivates on-device or nearby edge execution, where a single robot requires batch-1 inference outside the design point of LLM serving systems. Although VLA architectures combine familiar vision-language, autoregressive, and diffusion-style components, their runtime behavior in this batch-1 control setting remains uncharacterized. We characterize four representative VLA models on an edge GPU server and two onboard SoCs, using single-inference profiling and 43,200 closed-loop episodes. Action tensor dimensionality determines whether a stage is memory- or compute-bound, platform balance can shift that bottleneck, and GPU frequency scaling yields a platform-dependent energy-latency sweet spot. In closed-loop operation, overlapping inference with action execution creates an accuracy-speed-energy tradeoff, and no configuration is Pareto-dominant across deployment SLOs. These results guide joint design of VLA model architectures, hardware, and runtime policies.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
Embodied AI
Edge inference
Batch-1 serving
System characterization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action (VLA) models
Embodied AI
Batch-1 inference
Closed-loop characterization
Energy-latency tradeoff
S
Seonghun Jung
KAIST
S
Sieun Moon
KAIST
J
Jiyoung Jeong
KAIST
J
Jimin Lee
KAIST
Jaehyuk Huh
Jaehyuk Huh
KAIST
Computer ArchitectureOperating SystemsSystem Security