PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiency and fragmentation arising from repeatedly developing separate inference pipelines for physical AI models across cloud, edge, and endpoint deployments. To enable consistent deployment of Vision-Language-Action (VLA) and Whole-Body Action Mapping (WAM) models on single/multi-GPU systems and heterogeneous cloud-edge-endpoint environments, we propose PhyAI, a unified inference engine. PhyAI shares computation graph execution, memory management, and parallel serving infrastructure while preserving architecture-specific logic via model adapters. We introduce a novel control-time Roofline model to distinguish between compute-bound and environment-bound scenarios, guiding optimal execution strategies. Additionally, PhyAI integrates tensor parallelism, adaptive batching, and a custom caching mechanism to efficiently schedule Hopper GPUs. Experiments demonstrate 1.40×–4.65× speedups on models such as pi0 and GR00T; notably, Cosmos3-Nano-Policy-DROID achieves 1.18 s latency (2.08× faster) and 100 samples/s throughput (batch=32) on 8×H20 GPUs.
📝 Abstract
Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they often rely on separate inference programs. To unify them, we build PhyAI, a Physical AI inference engine with a single runtime that keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services. The same codebase runs vision-language-action (VLA) models and world-action models (WAMs) on single or multiple GPUs across onboard, edge, and cloud deployments. We used the adapter interface to add MiniCPM-Robot on the day of its release. PhyAI achieves 1.40x-4.65x speedups over the official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM-Robot. On Cosmos3-Nano-Policy-DROID it reduces latency from 2.46 to 1.18 s on eight H20 GPUs (CFG=2, TP=4), a 2.08x speedup. Specialized runtimes remain faster in several configurations, so our goal is one runtime with competitive latency rather than the fastest result in every case. Detailed profiles reveal why different models need different execution policies: on a Hopper-series GPU at batch size one, the pi0.5 action expert accounts for 8.8% of FLOPs but 57.2% of latency; at batch size 32 its share drops to 13.5% and throughput reaches about 100 samples/s. Cosmos3 remains generation-dominated and gains only 14.3% throughput as batch size increases from 1 to 16. We further introduce the control-time Roofline, which distinguishes inference-bound from environment-bound control; the measured pi0.5 points on four LIBERO suites are environment-bound while Cosmos3 stays inference-bound. Code and benchmarks: https://github.com/mingti-org/phyai.
Problem

Research questions and friction points this paper is trying to address.

Physical AI
inference engine
edge computing
cloud deployment
model serving
Innovation

Methods, ideas, or system contributions that make the work stand out.

Physical AI
unified inference engine
model adapter
edge-cloud deployment
control-time Roofline
🔎 Similar Papers
C
Chenghua Wang
Beijing University of Posts and Telecommunications
Daliang Xu
Daliang Xu
Peking university
mobile computingsystem software
D
Dongqi Cai
Nanjing University
D
Duojin Sun
Beijing University of Posts and Telecommunications
H
Hao Zhang
Beijing University of Posts and Telecommunications
H
Haoze Qian
Beijing University of Posts and Telecommunications
H
Huaiyuan Zhang
Beijing University of Posts and Telecommunications
J
Jinshuo Cui
Beijing University of Posts and Telecommunications
K
Kezhao Zhao
ModelBest
L
Longxi Gao
Beijing University of Posts and Telecommunications
Mengwei Xu
Mengwei Xu
Associate Professor, Beijing University of Posts and Telecommunications
Edge IntelligenceOperating System
R
Rongjie Yi
Beijing University of Posts and Telecommunications, MingTi Technology
T
Tianyue Zhang
Beijing University of Posts and Telecommunications
W
Weikai Xie
Beijing University of Posts and Telecommunications
X
Xiyuan Tan
Tsinghua University, ModelBest
Xuanzhe Liu
Xuanzhe Liu
Boya Distinguished Professor, Peking University, ACM Distinguished Scientist
Machine Learning SystemMobile Computing SystemServerless Computing
Y
Yingying Qin
Beijing University of Posts and Telecommunications
Y
Yiwen Lu
Beijing University of Posts and Telecommunications
Y
Yuan Yao
Tsinghua University, ModelBest
Y
Yuezhi Zu
Tsinghua University
Y
Yunhan Guo
Beijing University of Posts and Telecommunications
Z
Ziqi Guo
Beijing University of Posts and Telecommunications